V ========================== .... a
prevailing mis-
understanding about how hierarchical control works. It does not work
because the system has found a way to set the "right" references; it
works becuase the system has found a way to vary references so that
they end up producing the "right" perception -- the perception defined as
"right" by other references.
···
=================================
BN:
Let's take the reference signal for the perception of saying "no". This is
a single event perception. Also a syllable perception in parallel, as may
be seen if "no" is part of the event perception of saying "nobody". (The
same syllable perception in "notice" without the event perception for the
word "no".) Concurrently a category perception for the phoneme /n/ and
another for the phoneme /o/. (We'll pretend the diphthong is really just
one vowel, for simplicity.) The phoneme-category perceptions are inputs
to the syllable detector. The phoneme-category perceptions together with
the syllable perception are inputs to the "no" word-detector. Other lower-
level perceptions of nasalization (lowered velum), apical articulation
(tongue tip against upper alveolar ridge), tongue backing and raising with
lip rounding (for the /o/ vowel) are inputs to the phoneme detectors.
Acoustic and visual cueues are also inputs to the phoneme-category
detectors. Now, pick any one of these and there's not a whole lot of
variation that is possible. Less nasalization and you get "doe". Less
tongue raising and you get "gnaw". And so on.
++++++++++++++++++++++++++++++++
TB:
Examples like these occur reliably in discussions of differences between
speaking and tracking. (No putdown intended, only an observation.)
Each time they are raised, I try to understand the implication (sometimes
the claim) of great significance in the fact that changes in configurations
of the speech apparatus lead to changes in the sounds that are produced
and perceived. Every time, I come back to the thought that those
configurations vary any way necessary to produce perceptions (requested
by a higher level) of consequences other than the configurations and
sounds themselves. I still cannot see why the facts of configuration-
sound associations are different from, for example, the facts of various
configuration-movement associations we would observe were we to as
carefully monitor the activities of various motor units and joints in hands
and arms during arm waving, stick wiggling, gesturing and typing.
The constraint of "any way necessary" applies in either case; in neither
case does it imply "more than necessary," that is to say, "more than is
effective," with effective defined as production of the perception
requested in a higher order reference signal. Arm waving for its own
sake, as a perceived consequence of muscle actions, can be achieved by
many configurations of muscles and joints; waving an arm in the patterns
recognized by naval pilots as the signal to abort a landing on an aircraft
carrier is another matter. And it is another matter again for the stylized
hand movements of a Balinese dancer, or a Hindu practicing a mudra, or
a gang member in Los Angeles flashing a sequence of hand-signs. In
every case, a little more tension here or there, just a bit more pronation,
or extension, or rotation, and a whole other configuration occurs.
Whether the resulting configuration is significant to the actor or to any
particular observer, or class of observers, is yet another matter.
A personal example came to mind as I paused after typing those lines and
drummed my fingers on the computer table to accompany the Bach
"Brandenburg Concerti" playing on the radio. Those drummings,
contrasted with my now-resumed typing, contrasted with the movements
were I to type, not on my large keyboard, but on the tiny keys of the
Psion palmtop sitting on the shelf. Drumming can occur in many ways;
my movements when I use the big keyboard to type this message are
more constrained -- I can;t get too exc ited ort it fallls apart -- and my
readers will not understand. And if readers are to recognize what I type
on the Psion, I am even more severely constrained -- a tiny bit too much
extension of the right forefinger, not quite the porper arching, and I make
"u" not "j", and so it goes.
The same muscles and joints produce each of those results. In each
case, muscle activity varies, to produce a different intended result -- the
thing the actor is doing, not the actions seen by an observer.
+++++++++++++++++++++++++++++++++
BN:
So the word and syllable detectors specify the reference signals for the
phoneme detectors. The reference signal for producing /n/ is probably the
same for "no", "on", "paining", and "painting". That the behavioral
outputs differ in these four kinds of cases is probably a matter of the
physics of moving the tongue around rather than variation of the
reference signal.
++++++++++++++++++++++
TB:
Bill Powers already commented on this section.
[From Bill Powers (930525.1930)]
My point is that the reference signal that calls for producing the
perception of "no" is not itself the word "no." It is just a signal, which
specifies saying no not because of anything special about the reference
signal, but because it enters a certain comparator. Only its destination
makes it a "no" reference signal. And all that makes the corresponding
perceptual signal represent the word "no" is the way it is derived from
perceptual signals indicating that certain phonemes have been perceived.
The only way to get the "no" signal to appear is to supply signals from
phoneme detectors within a certain range of temporal and amplitude
limits. It's this concept that made Mark Olsen's remarks on "place"
coding ring a bell with me -- assuming this is anything like what he
meant. As you noticed.
The big objection to this whole picture is that this doesn't seem at all like
the way we experience "no" or phonemes or anything else. Some quality
seems to be missing. But what is missing, I think, is the context of all the
other perceptual signals that are in awareness at the same time.
..
Higher systems DO control their own perceptions by VARYING lower-level
perceptions (through manipulation of reference signals) -- but don't forget
that the higher level systems are perceiving different aspects of the
world. A system that tells a lower system to produce "no" does not
perceive the word "no". It does not know it is telling the lower system to
produce a word. Only the lower system can do that or know that; there
are no words at the higher levels. The higher level system perceives
something that is derived from the presence of the "no" signal and other
signals, and in this derived perception there is no longer any trace of
"no". To awareness, which can span several levels at a given time,
everything from the intensities and sensations on up coexists in the
experienced world, but a single level in the hierarchy always perceives
only its own specialized aspect of the situation.
+++++++++++++++++++++++++++++++
TB:
I second Bill's remarks, but he did not directly address the last part of
your paragraph, on the physics of moving the tongue. I am not sure I will
address it either, not in the way you might have intended, but it prompted
me to spend some time working my way through some of my
disorganized thoughts on that topic. Tongue wagging (talking) and eye
rolling (saccades, versions, etc) are often presented as special cases that
defy, or pose serious challenges to, the PCT model. (I am not saying that
*you* intended either of those, Bruce.) It has always seemed to me that
tongues and eyeballs DO enjoy rather special status: each is equipped
with an abundance of motor units and operates pretty much within the
protective confines of a body cavity or chamber (protruding eyeballs and
tongues are not all that common), relatively free from direct disturbances
(gravity excepted -- and generally overlooked by theorists). Those
circumstances, more than any other, probably contribute to easily-formed
impressions that feedforward, not feedback, controls the movements of
eyeballs, and that fixed commands (or reference signals) control the
movements of the speech apparatus. Those two systems operate in what
is probably as nearly an undisturbed environment as a human will ever
encounter -- precisely the kinds of conditions that allow a command-
driven model, like the one in "Models and Their Worlds," to seem
legitimate; and those are the conditions that lead Rick's "Blind Men" to
conclude that a control system is a command-driven (translate that as
'fixed-reference-signal') system. A few well-placed disturbances quickly
reveal that something else is going on. (Over a year ago, Gary Cziko
posted an ingenious demonstration of controlling speech for meaning, not
production, in spite of disturbances. What was that example?) People vary
their articulatory actions any wey necessary, the first time, when the
sensors for that system are blocked by local anesthesia administered by a
dentist, and they often produce intelligible speech (i.e., they create
a sufficient number of recognizible perceptions) even if they are chewing,
smoking, or (if they are PCT people who experiment on themselves) if they
hold their tongue against the roof or floor or side of the mouth, or against
the teeth, or between their teeth, when they hang upside down. And when the
disturbance is too much to work around, they no longer produce intelligible
speech.
Under typical conditions, the spectrograms of repeated utterances of a
particular speech sound are closely associated with a repeated productions
of similar articulatory movements, but it does not follow necessarily that
the movements result from a particular pattern of commands (as is sugested
in plan- or command-driven theories) or from invariant reference signals
(as in some interpretations of control theory).
+++++++++++++++++++
Bruce quotes Rick again:
(Rick)
++++++++++++++++++++++++++++++++
I (Rick) said:
What you (Gary)
describe is exactly what happens in Tom's cooperation experiment
You (Gary) say:
I don't understand how it is the same.
What I (Rick) meant was that a conversation is similar to what is going
on in Tom's experiment. In order to control a higher order variable (say,
"being polite") both people in the conversation have to be controlling for
a similar variable (like both subjects in Tom's experiment had to control
for "three lines in a row").
++++++++++++++++++++++++++++++++
BN:
You're jumping too far ahead again. In effect, you're changing the
subject. I'm not talking about cooperative control of conversational turns
or "being polite". What I (Bruce) said:
me (Wed 930519 12:26:11):
++++++++++++++++++++++++++++++++
The use of "s" vs. "th" to distinguish words (as in English sin vs. thin) or
of [y] vs. [u] to distinguish words (as in French tu vs. tout) is normative
and language-specific (culture-specific, group-specific), not universal. It
may follow from first principles of HPCT that apply to all humans (I
assume that it does), but if so it does so indirectly.
++++++++++++++++++++++++++++++++
So far as I (Bruce) can see, there is no parallel between Tom's experiment
and control of perceptions of contrast between words (or between parts
of words). You might get an analogy of some kind when B didn't
understand and A subsequently pronounces words more carefully, but
that's not the normal case that I wanted you to consider.
+++++++++++++++++++++++++++++++++
+
Rick replied:
[From Rick Marken (930524.1500)]
Whatever you say. I thought you were describing a situation where lower
level outputs had to be produced in a cooperative or coordinated manner;
Tom's experiment demonstrates one way of implementing cooperation
between control systems -- when these systems are housed in different
bodies. If you are not dealing with an example of cooperative control in
the control of contrasts between words then, fine, maybe Tom's
experiment is not a good parallel. But I'm pretty sure that it is.
++++++++++++++++++++++
TB (now):
(Whether the systems are housed in different bodies, or in one, makes no
difference.) I have the impression Bruce and Rick are talking about the
cooperation experiment at two very different levels. Not surprisingly,
Bruce seems to be keying in on my observations about the many ways
people varied their actions to communicate, while Rick seems to be
focusing on the modeling of interaction between two control systems.
Both levels are there to see. It is true, as Bruce says, that I did not (I
CANNOT) model the conversations, gestures and other means of
communicating. As things stand, the experiment only provides a setting
rich in actions by which participants intend to communicate.
I think Rick was pursuing the idea that cooperative interactions between
the models might offer a parallel to some aspects of speech production.
If he was, I agree. The two models produce the perceptual signals that
satisfy the (in this case, implicit) higher level reference of "three lines in
a specific configuration." I described "three lines in a row," but any
combination of two controllers could produce any of many possible
configurations. A change in the reference signal in one model, or in both,
leads to a new configuration. In a two-level system, a change in the
higher-level reference for a particular pattern would, if not matched by a
perceptual signal of that pattern, create error that acts as reference
signals for the two systems I modeled, and each would independently
produce one of the two two-line relationships that together result in
"three lines in the new configuration." (Hard to say; easy to do!) In that
case, the cooperative modeling illustrates how the same two independent
lower-level control loops can cooperatively produce any one of many
perceptual signals, corresponding to many different states of the
environment and to many differeny refeence signals from a higher level.
In none of those cases would the lower systems "know what they were
doing," in the sense of creating the three-line configuration. Each would
perceive its own little piece of a bigger pattern about which it was as
blissfully ignorant as Fowler and Bandura are about PCT. Were the model
fleshed out a bit, with a level of control loops below the ones I modeled,
the new loops could produce movements, with no knowledge of lines, or
of relationships between two lines, much less between three.
I imagine Rick was also thinking of how neither system controls its own
actions, which are controlled by the disturbance acting on the middle line
and -- in a way that still fascinates me -- by the actions of the other
system. Given a system with a reference signal for a particular
perception, anything that disturbs that perception "controls" the actions
of the system. I think those features of interactions between systems
might have implications for more elaborate HPCT models of speech
production.
++++++++++++++++++++++
Bruce commented on a post from me (TB):
(Tom Bourbon (930520.1410) ) --
BN said:
We are agreed that when members of a group of people share a setting
for a controlled perception, it is not because some superordinate social-
control system imposed a given setting upon them. Yet in the system
comprising the gather demo and its configurer/user, that is precisely what
is being modelled. The configurer/user imposes a setting for the proximity
perception on all the individuals in the model, from the outside. How
might this coming to agreement be put into the model?
+++++++++++++++++++
TB (now):
In the GATHER demo, the configurer/user "inserts" the reference signals,
and the gain factors as well. But when they do, I hope they do not
assume they can do the same things to living control systems. That is
the advantage of being a modeler/configurer; you have a chance to play
the part of a god. (I assume that only a god-like entity could so easily
tweak the reference signals of living systems.)
How might a coming to agreement be put into the model? If you mean
the appearances of coming to agreement, the characters in GATHER
sometimes seem to do that already; if you mean two or more systems
each acting to produce intended perceptions of coming to agreement,
they would need hierarchies deep enough to accommodate the setting of
reference signals for that level of perception. That is long way from the
ingenious programming of the multiple loops, all in parallel at the same
level, that Bill placed into the GATHERERS.
++++++++++++++++++
BN:
Again, you (Tom) describe how individuals in the gather model can be
given different settings from the outside, and how the
(Tom)
++++++++++++++++++++++++++++++++
results look sufficiently familiar that first- !time viewers of the program
characterize them as: "Like a (child-puppy) trying to be closer to (adults-
people) than the (adults-people) want them to be." Or, "Like a 'pushy'
person in a crowd."
++++++++++++++++++++++++++++++++
Again, the fact that these perceptions are not in the individuals in the
model means that relevant perceptions controlled by humans are not
being modelled.
++++++++++++++++
TB (now):
Of course that is the case. And it will always be the case, in modeling at
any conceivable level of complexity and detail. (That's why it is called
"modeling.")
In fact, a GATHERER controls only four or five modeled perceptions that
might be relevant to humans (or to birds, fish, dogs, apes, ants, robots
or any other social entities I can imagine). Amazing, isn't it, that an
observer could attribute so much knowledge and personality to a few
simple entities controlling so few relevant perceptions?
++++++++++++++++++++++++
BN:
Rightly or wrongly, people do have these perceptions and do attempt to
control them. Perhaps they are wrong and doomed to failure when they
attempt to control what after all are judgements and other stretched
analogies, and they would do better to control perceptions at a lower
level (as the Buddhists have been telling us). Nonetheless, that is what
people do, and we aim to model real people, I think.
++++++++++++++++++++
TB (now):
Of course people can have those perceptions. And, yes, sometimes they
are wrong and doomed to failure in their attempts to control them. Can
you suggest how we might model control of those perceptions? I don't
know how to do it. But I believe the primitive PCT modeling we *can*
do will be the substrate on which the really interesting models, like the
ones you request, will grow. Contemporary PCT models have scarcely
begun to swim in the primordial ooze.
++++++++++++++++++++
BN quotes and comments:
("Like a (child-puppy) trying to be closer to (adults-people) than the
(adults-people) want them to be." Or, "Like a 'pushy' person in a
crowd." Polite, offensive, pushy aggressive, nice, frenetic, sluggish,
patient, etc. As you (Tom) say, the observed agents do not have to have
these perceptions for human observers to perceive them as having them
and even as controlling them. The agents in gather are incapable of
having them. If they were capable of having such perceptions, would
they not probably perceive themselves as well as others in such terms?
+++++++++++++++++++
TB(now):
I think you are agreeing with what I said. Observers often see the
characters as "having" certain traits, qualities, or motives. I used that as
an example of how an observer can be wrong about what an actor-agent
is doing. But I also said that, given an HPCT control system of sufficient
complexity, those perceptions might occur within the system and could
become perceptions requested in reference signals.
++++++++++++++++++++++++++++
BN:
Let's look at speech vs. arm waving. I have the intention of
communicating to you that you are doing something that is not what I
wanted you to do. I have a range of means for doing this, including arm-
waving, words such as "no", and things in between like the vocal gesture
"uh-uh". If from this range I control a perception of saying "no" I don't
have a lot of freedom how to say it. You have to be able to recognize,
from whatever I utter, that I intend to say the word "no". Saying "oxi"
won't cut it for you (unless you know Greek), and anyway that would be
a different word-perception. What's involved in your recognizing that I
intend to say the word "no"? What are the settings that we must have
in common. (Remember, I have the intention of saying the word "no".
The intention to communicate denial is at a higher level, expressible by
various words and gestures.)
+++++++++++++++++++++
TB (now):
Agreed, Bruce. If you want to hear yourself say "no" you do not have
much freedom how you say it. In fact, you have NO freedom -- you must
do whatever is necessary to hear "no." So, too, in tracking, where it may
look to an observer as though I am free to move my arm and hand any
way I wish, but as soon as I decide to track, I am no longer free: I must
act to cancel the net disturbance acting on my controlled variable. And
once I grasp the handle in a particular way, or plant my elbow on the
table, or press my arm against my side, or hold a coffee cup in the same
hand as the handle ... . In every one of those cases, I give up even more
"freedom;" I put myself more certainly and rigidly (and vulnerably) under
the control of disturbances to my intended perception. I give up degrees
of freedom in my actions and am forced to use a smaller number of free
control loops, acting in more seemingly stereotypic ways. I have chosen
my perceptions, and I don't have a lot of freedom how to control them.
I do not mean to be flip or glib when I say this: The fact that a person's
choice of words is constrained does not seem to me to be different from
the constraints on every other kind of control behavior. And the fact that
some of the constraint seems to depend on the cultural or community
affiliations of the speaker confirms, for me, the idea that people (in the
universal sense) vary their actions (utterances) any way necessary to
achieve the control they are really doing -- which is hardly ever the
utterances we call speech. Most of the time. speech is like hand wiggling
and jabbing and turning; both are unintended means to intended ends.
(Babies, when they babble, seem to be among the very few who utter
speech sounds for their own sake.)
++++++++++++++++++++++
Bruce quotes:
(Tom)
++++++++++++++++++++++++++++++++ I
think the situation is exactly the same when I talk. While I am talking, I
am not controlling speech sounds. Exactly as was the case for tracking,
the reference signals for my actions vary, depending on error from above,
which depends on what I am really doing. I am not doing phonemes,
words, syntax and the like.
++++++++++++++++++++++++++++++++
I think you are confusing awareness with intention, i.e. control of
perceptions. My attention may be on talking a couple through the
cooperative task. Beneath that focus of attention I at a particular point
have the intention of uttering the word "no" in a way that the people I am
talking to recognize my intention of uttering that word, and behind that
(as well as behind concurrent gestures, etc.) to recognize my intention of
denying something indicated by other words and/or gestures. Beneath
the intention to utter "no" I have the intention to control the speech
sounds symbolized here /n/ and /o/. I am aware of the communicative
task. Beneath that awareness and in service of it I am indeed doing
phonemes, words, syntax, and all the rest.
++++++++++++++++++++++++
TB (now):
I don't think I was confusing awareness and intention. When I track, I
control (specify and hold invariant) my perceptions of the relationship I
intend to see. To control those perceptions of relationship, I cannot also
control (specify and hold invariant) all of the perceptions in control loops
below the level where the reference signal for the relationship is set.
Sometimes I can become aware of perceptions at those lower levels, but
they "just happen" while I control the perceived relationship. The
reference signals for the perceptual signals in lower levels cascade down
as errors from higher levels, but I do not intend them, nor am I aware of
most of them. I" do not intend all of the perceptual signals that occur
every moment, nor do (does?) "I" control the actions. I-the-intender can
remain serenely aloof while I-the-control-hierarchy control perceptions at
levels rarely know by the intender. Each of us inside here is ignorant of
the others and of what they are doing.
+++++++++++++++++++++++
BN quotes:
(Tom)
++++++++++++++++++++++++++++++++
The magic of words spoken to other people works through links in the
environment that are less tight than those when a person grabs a control
stick. Perhaps that difference in tightness has something to do with the
seemingly more constrained, but slowly-drifting, specificity of the sounds,
hence the actions, that work in a particular language community.
++++++++++++++++++++++++++++++++
BN:
Arm movements work to control the lines in any social environment. A
given sequence of utterances (in the manner of Chuck's note cards, say)
work in one social environment but not in another. Why is that?
+++++++++++++++++++++
TB (now):
Different intended perceptions.
+++++++++++++++++++++++++++
Until later,
Tom Bourbon