Percepts; reorganization; competing behaviors

[From Bill Powers (921006.1530)]

Bruce Nevin (921006.1321) --

But in your usage [of "percept"] it seems to evoke our subjective
experience ("qualia") rather than a neural current in a nerve
fiber:

The model is intended to assert an equivalence: that all objects of
awareness are neural currents in nerve fibers. These neural currents,
in turn, are representations, or in Wayne Hershberger's more
suggestive term, "realizations", or in drier modelling terms,
"functions" of an underlying order or Boss Reality knowable to us only
in the form of neural currents. Or, of course, neural currents derived
through more complex functions from the neural currents of lower
order. This includes all we are aware of, including our thoughts about
our experiences.

···

-------------------------------------------------------------------

The effort founders I think on the fact that you want to
communicate. In order to communicate, we assume prior agreement
about what appear to be CEVs in the world, and in the process of
communicating we verify and negotiate these and other agreements.

Yes. The only time it is profitable to keep the "pure" model in the
forefront is in the privacy of one's own investigations. Then it's an
essential aid in avoiding taking things for granted -- by exempting,
for example, the analysis one is currently developing from the status
of being "only neural currents." Communication between people who have
become easy with the basic assumptions of the model, however, takes on
a different flavor from ordinary communication that assumes a common
objective world. While it may seem that such people are talking about
controlling physical things and relationships in an external physical
world, the meanings being evoked in the participants include a
knowledge that the behaving system is really acting on its own
perceived world, and that the person describing the situation is
shaping this description of the behaving system with the understanding
that it is the describer's perceptions that are really being
indicated. All descriptions of this sort are simultaneously
theoretical propositions being tested for agreement with other
people's perceptions of the situation, with observations always
carrying the understood preface, "Here is how it looks to me, with
what evidence I have to justify my description:"
------------------------

But public stability is an assumption, a meta-agreement that is >part

and parcel of the "prior assumption" just mentioned. But this >is to
say only that we maintain reference signals for having >agreements
about CEVs in the public domain, and when discrepancies >result in
error signals we act to re-establish those agreements. >This seems to
me to be what compels negotiation and motivates the >search for
agreements. For the sake of the perception of public >stability,
these agreements once attained are attributed to the >world as being
knowledge of the world.

Again, yes. When we're communicating, the only sane thing to do is to
assume that our understanding of the agreement is the same as others'
understanding. But the CT-aware person never accepts agreement at face
value. As Martin Taylor might point out, and in other contexts has
done, there are far more degrees of freedom in the world than we have
under control, even during communication. What I agree to is a monster
object with more degrees of freedom than my means of expression have,
and I never know the extent to which your agreement refers to exactly
those degrees of freedom I am considering.

Some really nice statements in this post. Glad to see you getting up
to speed again.
-------------------------------------------------------------------
Martin Taylor (921006.1130) --

One place where I, and I suspect Greg, have not come to agreement
with you is in the limitation of the initiation of reorganization
to error in intrinsic variables--at least this is how I interpret
your term "critical error." You have, once or twice, allowed that
a form of critical error is some kind of integration of error over
all or part of the hierarchy, but for the most part I read you as
identifying critical error with errors in variables I would
characterize as body-integrity variables. My view of the mechanism
of reorganization is (I hope) the same as yours, but the
triggering, for me, is a sustained and particularly a growing error
in any ECS.

I don't want to put any a priori limits on what can amount to a
critical variable. I've proposed myself that error signals in the
hierarchy can amount to critical variables, in that minimizing them is
a goal that is compatible with the requirement that reorganization not
be dependent on the nature of the current environment (or acquired
knowledge of it). I'm sure that as we come to understand more about
how perception gets organized, and how levels of control develop, we
will find other such aspects of the hierarchy that are context-
independent and can qualify as critical variables of the inheritable
sort. Such things would certainly make the building of the hierarchy
more efficient than my simple overall error-driven random
reorganization.

I concur with your intuition about "a sustained and particularly a
growing error" as driving reorganization. My experiments with
reorganization, which continue sporadically, have convinced me that
there must be a strong rate-of-change component in the computation of
critical error signals. Just by trial and error, I found that the most
reliable method (so far) entails (roughly) making the rate of
reorganization depend on the rate of change of absolute critical error
TIMES the absolute critical error, so that the changes get smaller as
the remaining error diminishes. I realized only afterward that this
amounts to taking the derivative of the SQUARE of the error, which
neatly takes care of assuring positive values and that the result will
give the best least-squares fit to the optimum solution. You might try
this in addition to your (e*e + k*e*de)/G.
In other words, try d(e^2)/G.

Another tip. Rather than just make random corrections based on
critical error, I have found that it's best to randomly change a
parameter that determines the direction and speed of changes in
parameters. Between reorganizations, the changes then continue on each
iteration. This is strictly analogous to the E. coli method of
locomotion, in which, between tumbles, the organism continues to swim
in a straight line, continuously altering its relationship to a radial
gradient.

If there is an array of parameters d[i], I define an array delta[i].
On each iteration, d[i] has delta[i] added to it. This is like moving
in a straight line in hyperspace. As long as this movement continues
to reduce the measure of critical error by a sufficient amount,
reorganization is suppressed. Sooner or later, however, the movement
of the parameters in hyperspace will pass the point of closest
approach to producing zero critical error, and critical error will
begin to increase. Then reorganizations are commenced, which alter the
entries in the array delta[i] between positive and negative limits, at
random. Between reorganizations, the critical error is monitored, so
you can tell whether a reorganization left the error getting larger or
getting smaller. If it makes the error start getting smaller, you've
found a good direction in hyperspace (not necessarily the best, but
good is good enough). I haven't yet played with varying the number of
samples used to determine whether the error is getting smaller --
right now I accept any decrease as reason enough to stop reorganizing,
and any increase on a single iteration as reason enough to reorganize.
Refinements are obviously possible.

I have tried normalizing delta[i] so the sum of squares is 1. This
makes delta[i] into a unit vector. I think this works a little better
than just using raw deltas, but it may not be worth the computing
time, all things considered.

I still see some problems in switching signs, but I could be doing
something wrong. It shouldn't make as much difference as it does.
----------------------------------------------

My main difficulty with Hebbian learning is that it seems
unmotivated; that is, it's so local that I don't see how it could
be related to critical variables.

Let's take that in two parts. The motivation for Hebbian learning
is that it allows for gradient learning.

That's not exactly what I meant by "unmotivated." You actually
supplied the "motive" I meant when you said

The teacher decides beforehand which letter is to be output by each
of the 26 output units, and the idea is that if Q is presented, the
Q output unit should have a unity output and the others should have
a zero output.

Here, the teacher is supplying the missing motive for reorganization.
The teacher already knows that there is a "Q" present, and reorganizes
the network until it, too, reports a "Q". The critical variable is the
output of the network as perceived by the teacher. The critical
reference level is "Q". The critical error is in the teacher, who is
acting as a reorganizing system. The teacher's output acts to cause
reorganization in the network, which is terminated only when the
teacher experiences zero critical error. The things being reorganized
in the network have nothing to do in themselves with "Q"-ness --
they're just weights. The network itself doesn't care which output it
produces in the end. If the teacher wanted it to indicate "A" every
time there is a "Q" present, it could be made to do so (just relabel
the output lines).

So this is just like my model for reorganization, except that the
teacher is not inside the system but extraneous to it.

Your second example is about "teacherless" training:

In a teacherless system, the changes in weights are such as to
exaggerate differences among input patterns, leading units within a
level to decorrelate their outputs, or to converge to a common
output, depending on their initial sets of weights, the statistics
of the input patterns, and the local interconnections among units.

If there really are no criteria of error involved here, then such a
system will just converge to a state that expresses its properties,
which are fixed. This would not fit my idea of reorganization; it's
more like a fixed algorithm. I would guess that it's less capable of
learning than the kind with a teacher. But maybe this is an important
kind of network anyway, in that it would have many possible output
states and could provide ABITRARY discriminations, for acceptance or
rejection by a reorganizing system. I suspect, however, that the
meanings of the discriminations are subject to a lot of interpretation
by the human beings who are looking at the results. As you've
described this sort of system, I can't see anything that would
constrain it to making USEFUL discriminations. Aimless complexity
isn't necessarily useful.

I think that in the overall picture, a teacher is necessary. But the
criteria this teacher should use must be relevant to the system
getting reorganized, not placed externally in a different organism. My
reorganizing system is a teacher that uses criteria relating to the
functioning of the system itself: the teacher wants all the control
systems together to have as little total error as possible.

In the Little Baby project, the second (probably) experiment will >be

to start with a random set of output connections, and use >Hebbian
learning on the perceptual input functions to see whether >the baby
can learn to perceive the world in a way compatible with >its (fixed
randomly assigned) outputs.

I can tell you already that this will work, using my method of
reorganizing outlined above, with up to 20 independent control systems
controlling through 20 shared environmental variables. The critical
variable is total squared error across all systems. Reorganization is
applied globally, not based on each system's error.

I didn't have the patience to let 50 systems converge, but it looked
as if they were headed that way (that's 2500 weights being
reorganized). There's no reason it shouldn't work with any number of
systems. When the output connections have fixed signs (although chosen
at random), and you just reorganizing the input weights (n weights per
system when there are n systems controlling n environmental
variables), convergence always occurs and with reasonable efficiency.
And this is the worst case, because there are no spare degrees of
freedom. Of course it's possible that the choices of output weights
will preclude a solution (I think). I haven't run into that yet,
although some choices make convergence definitely slower.

The third, and most interesting experiment, if the first two
succeed, is to see whether reorganization and Hebbian learning can
be used together, actions adapting to perceptual input functions
that change toward more controllable forms. That's a little way
off yet, but I could imagine that at least some results from one or
two experiments might provide a Christmas present for the group.

That would be a nice present.
---------------------------------------------------------------

RE: competing behaviors

When one of the mutually competing behaviours acts so as to provide
a percept that satisfies the higher-level ECS's reference signal
(taking the car gets me closer to my destination), the reference
signals for the other would-be competitors are reduced.

What reduces them? I still don't see how you can be riding a bicycle
and driving a car at the same time.

Then the higher-level reference is not being satisfied, and as you
say, the output of the "taking bike" system increases, to the point
where it overtakes the "taking car" system and inhibits it.

Where are these outputs acting, outside the system? What kind of
outputs are they? Come on, Martin, you're waving your arms.
------------------------------

The world is unstable, with unpredictable disturbances. Sometimes
pushing on something that always worked one way now makes it move
the other way. There are disturbances and conflicts that can cause
perverse effects even if the world actually is working the way you
expect.

This was more a defensive paragraph, against those who would
correctly point out that no control system always finds its
percepts moving in the way that they usually do, given a particular
output signal. If they did, you might as well have a simple pre-
planning system of the type you call "cognitive". Control systems
exists as a defence against the unpredictability of the world.

The levels exist to eliminate the need for reorganization or random
processes. If control of certain things often entails switches of
sign, a higher level system monitoring the relationship between
direction of action and direction of effect will be acquired to make
the required switch without any trial and error, immediately (where
"immediately" means, apparently, in about 0.4 sec, according to Rick's
experiments). Ordinary variations in disturbance or parameters do not
require any adaptation from a control system; it just acts as it
usually acts.
-------------------------------------------------------------------

The only way for an error to be small is for the output to be in
the state that brings the perceptual signal nearly into a match
with the reference signal and keeps it there.

I disagree. There are lots of situations in which error is small
by chance.

Nope. If external forces bring the error to zero by chance, the output
of the system will drop to zero. If it doesn't, it will CREATE an
error. The output is always in the state that brings the perceptual
signal to a match with the reference signal and keeps it there, even
when that required state of the output happens to be zero. "Chance"
has no effect on error signals. Error signals are affected by physical
variables: outputs and disturbances.

Anyway, as I said, "probably" was for accuracy rather than as an
assertion that the probability was far from unity.

That makes quite a difference.
----------------------------------------------------------------
Best to all,

Bill P.

[Martin Taylor 921008 20:15]
(Bill Powers 921006.1530)

I'm delayed in this response, and I think that I may not be able to respond
much over the next 3 or 4 weeks, first because of other pressures and then
because I'll be away. I owe Rick a response on statistics based on a private
interchange he suggested be continued publicly, and I hope I'll get to that.
But here's a little something, anyway.

RE: competing behaviors

When one of the mutually competing behaviours acts so as to provide
a percept that satisfies the higher-level ECS's reference signal
(taking the car gets me closer to my destination), the reference
signals for the other would-be competitors are reduced.

What reduces them? I still don't see how you can be riding a bicycle
and driving a car at the same time.

If the reference for "see myself as riding a bike" is a result of the
output from the ECS that is controlling for "see myself arriving at the
destination," then nearing the destination will reduce that reference.
I suspect, however, that since we are working above the category level,
there are threshold effects as well. Be that as it may, when I am at the
destination, there is no error in not seeing myself taking the bike.

Then the higher-level reference is not being satisfied, and as you
say, the output of the "taking bike" system increases, to the point
where it overtakes the "taking car" system and inhibits it.

Where are these outputs acting, outside the system? What kind of
outputs are they? Come on, Martin, you're waving your arms.

Well, I can't do that and type at the same time, can I?

In the scenario I painted, the output for taking the car was indeed acting.
It was acting on the CEV that we call driving the car. But the world made
that CEV immovable (the car didn't work, so it, too, was immovable). The
output of taking the bike was not capable of causing overt actions in the
real world so long as it needed ECSs at lower levels that were acting on behalf
of the "taking car" ECS. But when its output became large enough to cause
a switch (which I think must occur at the category level and above, there
being no intermediate possible reference levels between categories), then
those lower ECSs would cease supporting "taking car" and start responding to
reference levels that ultimately derived from "taking bike." This doesn't
imply internal inhibition. It's just that the relevant CEVs do not admit
to simultaneous control in the real world, because in the real world their
control requires some lower-level CEVs to be controlled to more than level
at the same time, which can't happen.

I don't perceive arm-waving, and I'm certainly not controlling for perceiving
it. The linkages and interactions seem quite clear an unobjectionable to me,
and I'm not quite sure what it is you are having difficulty with.

I do have to admit that I have not built a simulation.

Martin