[From Bill Powers (941201.0950 MST)]
Hans Blom (941130a)--
Whereas you are more interested in control, I am more interested in
learning/optimization/reorganization, its laws and its possible
implementations.
Fine, there's room for everyone.
Isn't it obvious that learning has to produce (better and better)
control? What we learn through the process of learning is how to deal
better with the world around us, i.e. how to better (with more
precision and/or more efficiency) achieve our goals and, possibly, how
to discover more worthwhile goals. Is any discussion about this topic
still necessary?
Among the converted, you mean? Possibly not. But very little of the
world of the life sciences is converted at this point. And even on the
subject of control, there isn't complete agreement among those who say
they believe in it on what is and is not control.
Why is obtaining a food pellet a "good" thing, where is this goodness
defined, and how does the goodness or badness relate to altering the
organization of behavior?
Where do you see the "good" and where the "bad"? Certainly not in the
outside world. Maybe in your gut feelings, but those may be difficult
to communicate to someone who has very different ones.
I see it as defined within the organism: anything for which there is a
built-in intrinsic reference level significantly greater than zero is
the definition of "good" for that organism; anything with a zero
reference level is "bad." That is, the value of all experiences (as far
as the learning of behaviors is concerned) is determined by how they
impact on these fundamental variables.
I am more concerned with learning as a process: how come that you
gradually (or suddenly) do things better than before? What has changed?
What has been discovered?
Yes, that is the question to be answered. I maintain that you can't say
what had changed or been discovered until you can specify how the system
was working both before and after the change. That means modeling
behavior -- performance -- to the point where you can state measurable
parameters. Once you can do that, you can point to what has changed as a
change in parameters.
That does not mean that I find your questions uninteresting. It is just
that I do not believe in deep answers to "why" questions; such an
answer always points to and is an analogy with something that we know
already, i.e. it is a (newly discovered) correlation with something
that exists already.
When I ask "why" I am thinking in somewhat different terms. In a spinal
reflex, we can show that there are reference signals which determine the
state to which the stretch-tension control system will bring its
perceptual inputs. But why that state? The answer can be found only in
the operation of a higher system that is setting the reference signals
as a means of controlling some higher-level perception. The answer to
"why" at one level is "how" at a higher level. How do we raise an arm to
point to something? By varying reference signals for the spinal control
loops. Why are those reference signals varied as they are? In this case,
in order to point to something.
I also do not believe in models as having explanatory power. A
simulation is, basically, a trial-and-error test which generates
outcomes based upon a selected number of hypotheses and on assumptions
about how to connect them.
A model which is strictly ad-hoc has explanatory power only by luck. But
all models have structure and properties, which we can treat as
propositions about how the physical system is actually organized. In the
1973 book I did my best to find evidence in the nervous system relevant
to the structure and properties of the model, and what I found on the
way to doing this had a lot to do with the ultimate form of the model as
presented.
Another thing models and more specifically simulations do for us is to
demonstrate unexpected consequences of our own assumptions. I would
never have guessed at the basic properties of control systems if I
hadn't studied them in simulation and found the counterintuitive
relationships that emerge from a closed-loop system. Models which are
presented _without_ having been tested in simulation often don't work
because of subtle internal inconsistencies or because what seems a
complete statement turns out to be ambiguous or incomplete when actually
reduced to a simulation. Simulation is especially important where the
system is so complex or nonlinear that there are no analytical
techniques for solving the system equations. Without simulation, one can
only try to find some approximation that can be handled analytically,
which poses the danger that one will end up analyzing an imaginary
system that is importantly different from the real one.
If those outcomes correlate well with something that we know already,
we tend to say that we have an "explanation". In my opinion, however,
we have only discovered a new analogy.
I think your view of simulations is too limited. I have discovered many
things from simulations that I didn't know already. There is more to a
simulation than simply reproducing a behavior. Simulations are made of
components each of which is an hypothesis in its own right, subject to
test against the real system. A simulation proposes a relationship among
components, and the relationships are just as important as the behavior
that results from them.
Moreover, the number of models that may "explain" certain outcomes is
infinite. How can we ever decide which is the "best" one, except in
subjective terms of "elegance" (as mathematicians do in their proofs),
"conciseness" (number of symbols) or some such.
In the first place, not all models behave as their authors claim they do
(when expressed as a working simulation). In the second place, a good
model is built of components which have actually or potentially
verifiable existence. The point of a simulation is often to reveal what
will really happen when known components are connected together in known
ways (as opposed to what someone may claim will happen). In my view, the
best model is one in which every component and every pathway can be put
into correspondence with a real component and pathway in the real
system. To the extent that we can do this, the model explains what that
arrangement DOES. Even when we can't do this, my view is that we should
build our models so they could be tested against the real system when we
become able to identify comnponents and pathways. Every parameter, for
instance, should have at least potential physical meaning in the system.
I don't know whether you are familiar with the ALife literature, but
the simulations show some surprisingly familiar results using only very
few and simple hypotheses. The behavior of a chimp male, for example,
can be "explained" (to an 80% fit or so, the authors say) by the
following simple "laws": "if you see another chimp, go find out if it
is a willing female; if so, mount her; if not, go look for food". Quite
simplistic, but also quite convincing that a few simple hypotheses can
"explain" seemingly very complex behavior.
Yes, that's another advantage of simulations. What seem to be complex
behaviors often turn out to emerge quite naturally from a system with a
basically simple structure. The Crowd program is an example: a
collection of individuals with at most three active control systems in
each one produces mass phenomena similar to real ones, without the need
to invoke any complex "social" laws.
Simulation is basically a form of mathematics. Take Euclidean geometry:
Choose a few axioms (which has, what is it, 5 or 6 axioms?), and devote
the remainder of your life to find out the implications of those
axioms, i.e. develop more and more theorems.
That's a rather dismissive view, isn't it? It implies that mathematics
is a waste of time.
That does not mean that I do not use simulations. They are valuable in
the discovery what the actual behavior of a designed system is, whether
that behavior fits the system's specifications, and whether the
behavior is "surprising" (unexpected) under certain conditions.
Yes, indeed, my view exactly. Even if you are simulating a system that
you understand (you think) completely, running the simulation and
playing with its parameters will teach you things you didn't know
before.
The type of reverse- engineering that is your objective is awfully
difficult if one has to treat the system as a black box and is not
allowed to investigate the insides.
Who says we're not allowed to investigate the insides? Perhaps we,
personally, have no license to do this, but there are plenty of people
who do, and they publish what they learn, and we can take what they
publish into account to the extent possible.
The engineering problem is this: given the specifications of a system
(however formulated), an infinite number of designs is possible. So
which design to choose? In practice, the design seems always based on
the methods that the engineer is most familiar with...
In Little Man v. 2, the basic design of the control systems was a direct
translation of the functions and pathways known to exist -- the
neuroanatomy -- simplified somewhat for purposes of modeling. So a heavy
constraint was put on possible forms of the model. The model not only
moves the arm, but it does so using components and connections that are
known to exist. This is not just curve-fitting. In fact, I was much
impressed by the compactness of nature's design and by the way it
achieved dynamic stability in what seems to me the simplest possible way
-- which I would never have thought of. I was prepared to add all sorts
of filtering and stabilizing circuits to get it to work, but all I did
was reproduce the basic arrangement known to exist, and the system was
stable from the start. I think you're talking about someone else's way
of modeling.
···
-----------------------
... reorganization, of which there seem to be two types: hill-climbing,
as in your ecoli simulation, which has a built-in method of how to
fine-tune parameters (as Bruce notes, this is better not called
learning but hierarchical control); and trial-and-error learning, as in
Bruce's ecoli simulation, where a built-in hill-climbing procedure does
not exist ...
I like this distinction of yours -- a few days ago you expressed it as
gradient-climbing and discovery learning. Actually, Bruce's model is a
gradient-climbing model, because it is organized around an already-
existing control system with a single adjustable parameter that
determines the delay between tumbles. The model does not have to
discover that the time-rate-of change of nutrient can be affected by
adjusting the timing of random tumbles -- or for that matter, that
controlling the rate of change would be a good idea. All that is given
in the basic structure. The only thing that remains to be learned is how
to adjust the parameter to produce the desired result.
But there seems to be a third type. Let me give you an example. For
many years I rode my bicycle to the university along the same route,
until once I met and started to accompany a colleage, also on bike, who
led the way ALONG A DIFFERENT, SHORTER ROUTE. After that, I started to
travel the newly discovered way. Did I have an error before? Not that I
know of.
If you had no error before, you weren't controlling anything before. All
control systems have errors, all of the time save in instants of passing
through zero error. I think your assumption that the newer shorter route
reduced no errors for you is too hasty. Was there no cost in traveling
the former route?
However, you have a point in that the acquisition of knowledge is a form
of learning, and sometimes new knowledge can lead to better forms of
control without reorganization. If I see a man jumping up and down on a
wrench and cursing because the nut won't turn, it might be that I know
that this piece of machinery happens to use left-hand threads, and save
the man a great deal of sweat simply by passing this fact on to him
verbally. All he has to do is grasp my meaning, understand the
implications, and reset his reference signals accordingly: the nut will
then turn, and we can say that the man has learned how to loosen a nut
on that kind of machinery.
Of course the contingency acquires no new properties. I thought I men-
tioned that. Of course it was the _cat_ that changed. We do not
disagree at all, although you seem to think so. But you do not go far
enough. There are no "causal relationships" in the physical world; they
are internal in us. Why is _this_ so hard to see?
Come on, now you're just looking for revenge. In the context of
practical epistemology, it's just a mistake to say that a consequence
can select a behavior. Of course you can always flee to a higher level
of abstraction and reply that it's a mistake to say anything, but my
point was much too simple to warrant such a strategy.
What about reference signals? AAARRRRGH!
How about the above example? Did I have a reference signal for
following a shorter route to the university?
You did after you adopted it. But the change in route was a change in
behavior, and I would presume it occurred because its _effect_ was to
reduce error in the control of some variable of importance to you, such
as how sweaty you were when you arrived at work, or how tired, or how
late. We do not control behaviors; we control consequences of behaviors.
Our reference levels do not specify acts, but outcomes.
Yet the perceived consequence of an action led me to change my
behavior.
That consequence didn't occur until after the behavior changed. And the
behavior was not the consequence; the consequence was whatever of
importance to you changed as a result of changing your behavior. If the
consequence had not been closer to what you prefer, the very same
consequence would have led to abandoning the behavior that caused it. It
is not the consequence that determines whether you retain the behavior;
it is the difference between the consequence and the consequence you
want. That is why I asked, "what about the reference signal?"
Some learning seems not to depend on reference levels that are a priori
given, but leads to the installation of a new reference level (or a new
control system?) _after_ a discovery.
All this will be clearer if you think in terms of levels of control. The
reference signal for which route you take has to be changed _before_ you
change routes -- otherwise you would go by the old route. Once you have
changed routes, that control system has done its job: made the actual
route match the reference route. A consequence of this may have the
effect of reducing error in the system that chooses reference signals
for routes; if so, we have to look at the higher system for the
explanation, not the system that executes the movements involved in
following a specified route.
This example may illuminate your concept of "discovery" learning. In
choosing a new route, you are not learning any skill you didn't have
before, nor are you fine-tuning a skill. Instead, you're selecting among
different ways of using existing skills, at a level higher than the
skills. We could think of a collection of routes about which you know,
only one of which, of course, can actually be employed on any one trip.
So the higher system, concerned with controlling things that are
affected by which route you take, has to pick one reference-route for
the lower systems to execute. The point is not to select a route, but to
control something that is affected by the choice of route, such as how
long the trip takes. And barring knowledge of all possible and feasible
routes, the only way accomplish the result is to try different routes
and see what the result is. So this is an example of learning in which
the mechanism involves trial-and-error selection among different lower-
level behaviors, each of which is a well-learned control behavior in
itself. The higher-level system is reorganized to send its output signal
to one subsystem rather than to another -- there is some sort of
switching of the output signals.
-----------------------------------------------------------------------
Bruce Abbott (941130.1915 EST)--
Although my background is strongly Skinnerian, my views are not. (I've
cited Skinner more often in CSG-L than previously in my entire life.)
Yes, I'm gradually understanding where you're coming from.
I'm more interested in attempting to apply and evaluate PCT in a field
traditionally dominated by reinforcement theory than in playing the
role of Resident Behaviorist. Partly this means rethinking many issues
and concepts from a PCT perspective that are dealt with in very
different ways from a TRT viewpoint.
That puts us on the same side. But your role as Resident Behaviorist is
exceedingly important, because there must have been a time when you
found the Skinnerian framework convincing, and knew no other. So you
have gone from there to here and know the route (whether you went by
bicycle, like Hans, or not). To me, this means that you should be able
to pick out milepost experiments that can lead other behaviorists to
grasp PCT and see that its interpretations make sense. You have standing
in that community, and your publications will be accepted. As you can
see, my attitudes are purely selfish.
My "defense" of the law of effect is one result: I recognized a
principle at work that did not seem (to me at least) to contradict PCT,
although it clearly needed to be reformulated from that perspective.
In so doing I tried to differentiate between the empirical law of
effect, which is just a description of what one can see happening, and
the theoretical law of effect, which includes unobserved entities like
associative connections. I have been arguing for the former and not
the latter.
I see and accept that. You must forgive the prickliness with which
behavioristic concepts are greeted on this net -- we may be pioneers in
some regards, but we're clearly not diplomats. The biggest problem is in
sustaining the long-term view. This has been a long, frustrating, uphill
battle, and there's always the temptation of wanting it to be over while
we're still around to enjoy the result. But the end will not come
suddenly; it will come through the building of one small advance on the
ones before, and the battle may go on for a long time yet. One person
here, another there, joins in, and each becomes a center from which the
basic idea can spread to new places. That's how PCT will eventually
become accepted.
I agree that Thorndyke was poised at the right threshold. So, in a way,
was Tolman, with his "purposive behaviorism." And so was Skinner, in the
sense that he set up the first real control-system experiments and tried
to make a break with simpler concepts of behavior. One day someone will
go back over all these historical developments and ask the question,
"What kept them from taking the next step?" I think we'll learn
something about how science progresses from that. Psychology has been
bumping up against the same barrier for close to a century.
Thorndike's law of effect, then, asserts that those activities taking
place within a given situation will tend to be repeated if they bring
about a satisfying state of affairs or prevent or eliminate an annoying
state of affairs.
As you describe it, a perfectly acceptable description -- of the
problem. What's unfortunate is that so many psychologists took it as a
solution.
From this mix would be selected those behaviors that proved effective
in reducing those errors in the cat's perceptual control systems that
had been induced by the disturbances of confinement and lack of access
to the tasty treats provided outside the box.
Yes. I hope we can some day find a coherent way to tackle a problem that
is implicit in all these descriptions -- the failure to think of
behavior in any terms but events. "A behavior" creates "a consequence."
But behavior is continuously variable, and is organized at more than one
level. "A behavior" is seldom effective in reducing an error: what is
effective is for the amount and direction of behavior to be properly
related to the amount and direction of error. A lot of psychology is
oriented toward seeing behavior as a succession of point-events: this
happens, and because of it that happens. This is like considering
falling bodies only in terms of whether they fall or don't fall, leaving
out position, velocity, and acceleration. Behavior is a continuous
ongoing process, not a succession of occurrances. This was one of
Skinner's main contributions -- to see in bar-pressing and
reinforcements not just a series of discrete events, but _rates_, which
are continuously variable. He put one foot into the world of continuous
variables, but only one.
I agree with you that one should strive whenever possible to provide
mechanisms whose proposed components might be found to be represented
in some real form within the actual organism being modeled, as in the
original e. coli simulation. It is for this reason that the tests
between my specific implementation of a "law of effect" model and the
"PCT" model seem to me to be beside the point. On the other hand, they
do very On the other hand, they do very nicely illustrate how one goes
about evaluating the relative merits of competing models.
Yeah, let's go on to other things. I was going to develop a reorganizing
model of E. coli, but it's really more trouble than it's worth: putting
a reorganization model on top of a behavior that's already basically a
reorganizing process! I've been thinking of going to a logarithmic
perception of concentration to cut down the dynamic range, but then I
think why am I doing this? I'm just getting bogged down in the details
and we already know it can be made to work, with enough effort. That
random output just makes everything more murky to talk about -- let's
get on to what we really want to model.
But it has, as you say, been a nice illustration of how arguments go
when disciplined by the requirement of producing a working simulation.
Think of it: if all theoretical arguments had to produce simulations,
how thin the journals would be!
And I think that non-modelers will be able to follow models of operant
conditioning much more easily than the stuff we've been going through.
I do think we should take a little more trouble to explain the meaning
of these models, and describe their behavior, for people like Chuck
Tucker who has asked for this, and Mary, too, and probably quite a few
others. Maybe whenever we arrive at a concensus about a model at some
stage of development, we should fix it up in user-friendly form and post
the runnable version on Bill Silvert's server for people to download and
try, with a decent writeup.
An analysis that deals only with the steady state and is silent on how
that state was achieved will not successfully compete with one that
appears to do both (even if appearances are deceiving.)
We can do this to a degree with hill-climing learning, but it's going to
be tough to handle discovery learning (in Hans Blom's terms). I doubt
that we'll be able to model the process by which the rat discovers,
while nosing around in the cage, that pressing a bar causes a piece of
food to drop. Well, maybe some day, but I sure don't feel ready for that
now.
I think that the most useful immediate result we can get is to show that
a very simple hierarchical model of _fixed_ organization can account for
behavior over a wide range of schedules, and maybe over types of
schedules as well. What this will do is to show that changes in behavior
that had been thought of in terms of learning can be handled in terms of
the performance of the right kind of system. I'd like to see us at least
try something in this direction. Everything we can account for in terms
of a performance model is removed from what we have to account for with
a learning model, leaving the problem of learning simpler than before. I
don't know that we can really do this, although I had some success with
the Staddon/Motherall data. But it seems worth a try, considering the
payoff if it works.
As a start in this direction, how do you define learning, of the type
seen as an organism constructs an appropriate behavioral output sub-
system that brings a perceptual variable important to it under control?
It is only by having an explicit definition that we can agree whether a
particular model does or does not learn.
Right on. If you allow higher levels of control to operate through
altering parameters of lower systems as well as their reference signals,
the whole idea of "learning" begins to look different. Some time ago I
would have defined learning as a change in the parameters of the system,
but if there are systematic changes in parameters, brought about by a
control system of fixed organization, that idea sort of loses its force.
I think the ultimate definition will have to do with acquiring abilities
that were not present AT ALL to begin with. But I guess I would really
like to see how much we can account for, that looks like learning,
without actually invoking that concept. As we do this, the meaning of
learning will probably become clearer.
-----------------------------------------------------------------------
Best to all,
Bill P.