PCT and TRT

[From Bill Powers (941129.0600 MST)]

Bruce Abbott (9411various) --

While we've been up to our elbows in details of modeling probabilities
and contingencies, we've sort of lost track of the big picture. It's
been interesting to work with E. coli, but a microbe that uses random
tumbling as a means of steering isn't the best example for seeing the
regularities -- and lack thereof -- in behavior.

The real problem with the concept of consequences selecting the behavior
that leads to them is that in general there is no one behavior that will
lead to them. Consequences are a joint function of behavior and
independent variations in the environment. The driver, the road, the
wind, and the car all contribute about equally to the forces that steer
the car. The consequence of keeping the car in its lane couldn't
"select" a steering action that will have that result, because there is
no such steering action. This simple fact, which is true of most
behavior, has been overlooked in our conversations lately.

The appearance that the same behavior produces the same consequence
results from careless observation. Skinner was not a careless observer:
he recognized the problem here, although he didn't have the answer. He
rejected stimulus-response theory because he saw that neither stimuli
nor responses repeat from one instance of a behavior to the next. The
nearest he could come to a solution was to state the problem: behavior
has to be defined in terms of _classes_ of actions, and those classes
can be defined only in terms of consequences. "Bar-pressing" behavior is
actually a large class of actions -- infinite, if you look closely
enough and think about it -- which are held together in our minds only
because they have a common consequence: the bar gets depressed. Bar-
pressing is not really an action. It's an outcome of actions, and the
real actions can vary all over the place.

That way of putting it doesn't reveal the real nature of the problem,
because it focuses on different actions that happen to have a similar
consequence. The rat can use its nose, paws, mouth, or rump to depress a
bar and can do so from an infinity of different postures and
orientations relative to the bar. But if this is all one gets from
Skinner's observation, the main point has been missed.

The main point is that in general it is _necessary_ to use different
behaviors if the same consequence is to be brought about. If a pigeon
happens to find itself to the right of the key when the discriminating
stimulus light comes on, its response must be to move to its left if
it's to peck on the key. Yet a few moments later it may find itself to
the left of the key, or in front of it, or ten inches away from the key,
or with its beak poised exactly over the key. Somehow the same
consequence, the application of a peck to the key, must select
_different_ actions by the organism under different conditions.

The discriminative stimulus seems to be one answer to that problem, if
we recognize that this stimulus is a perception experienced from the
point of view of the organism (and not an objective state of affairs).
Now we can see that each different starting situation amounts to a
different set of discriminative stimuli, which can lead to a whole
collection of different behaviors -- just the behaviors needed to
compensate for the changes signalled by the discriminative stimuli as
experienced from the point of view of the organism. All that's necessary
is for the organism to discover what action should go with each
different discriminative stimulus.

But now we come up against the core of the problem: there are few
situations in which discriminative stimuli exist for each possible
environmental disturbance that can alter the relationship between
behavior and a given consequence. In many cases where such stimuli seem
to exist, there's no way they can be quantitative enough to account for
the quantitative precision of the consequence. And consequences are
often repeated under novel conditions, where even if discriminative
stimuli of sufficient quantitative properties existed, there has never
been a previous experience with them to permit acquiring the necessary
instrumental behavior.

If every possible behavior that could lead to a given consequence were
signalled in a quantitative way by some environmental indicator, and if
the organism had experimented long enough to have acquired a three-term
contingency for each of millions of different situations, then the
concept of consequences selecting behavior might stand up in court. But
those prerequisites don't exist under most real circumstances.

The control experiments in the paper by Bourbon and me, _Models and
their Worlds_, were designed to violate, one at a time, the assumptions
behind the explanations above, and to show that consequences were,
nevertheless, under control by the organism.

Consider the pursuit tracking experiment. A target moves irregularly on
the screen, and the subject uses a control handle (or a mouse) to keep
the cursor as near to the target as possible.

The explanation offered by most behaviorists for this behavior is that
there is a visual stimulus on the screen, a consequence of behavior,
that selects the handle movements which produce that consequence.
Eventually, this consequence (also playing the role of a discriminative
stimulus) is usually defined as the relationship between the cursor and
the target. If the cursor is to the left of the target, the required
behavior is to move the cursor to the right, and so forth. This seems to
fit the behavior very well.

But then we introduce a disturbance which adds to the effect of the
handle position on the cursor position. This disturbance is derived from
a random number generator, successive values being smoothed to limit the
bandwidth to about 0.1 Hz. There is no indication on the screen of the
magnitude of the disturbance. All that the subject knows is that the
cursor no longer faithfully follows the movements of the handle. Even
with the handle held still, the cursor wanders in a random pattern on
the screen.

Now there is no way for the consequence of the behavior, the cursor
remaining near the target, to select the pattern of behavior that will
create that consequence. Sometimes when the cursor is left of the
target, the subject is moving the handle to the left. There is no
discriminative stimulus to indicate what behavior is appropriate. The
random disturbance pattern never repeats, over hundreds of experimental
runs, so there is no pattern to learn. Yet what we see is that the
cursor follows the target very closely on the screen. Because the cursor
follows the target so closely, there is nothing in its pattern of
movement to indicate the magnitude of or pattern of the disturbance,
which follows an entirely different pattern. Yet we see the handle
movements reflecting the magnitude of the disturbance at every moment,
plus enough more movement to keep the cursor near the target.

I would like to see us turn from E. coli to these tracking experiments.
Behaviorists have offered explanations of them before, but you are the
very first behaviorist I have encountered in 20 years of arguing with
them who understood anything about simulations, and if behaviorism has
any defense against the phenomena to be found in these experiments, you
are the one to find it (you and Sam Saunders, who seems to have
disappeared). We have over a month before your animal studies can
commence. I'd like to see us spend at least a week or two on the
tracking experiments, to see whether any competing model can reproduce
the results as well as the control model does. If there is any way to
refute the succession of arguments above, we should find out about it.

···

----------------------------------------------------------------------
Best to all,

Bill P.

[From Bill Powers (941204.1020 MST)]

Bruce Abbott (941202.2000 EST)--

Some while back I posted ECOLI6a, a version of my nutrient-regulating
e. coli that included Bill Powers's FramePlt graph plotting routine to
make visible the changes in dNut and nutrient levels ("fuel") over
time. I never heard a word from anyone about this--did anyone try it?

I was sure I had commented. I just went back and looked at it, and found
that it uses a timer to vary the tumbling frequency -- it looks as if I
may have modified the code in playing around with it. Could you send me
the original again (direct), so I can make sure I'm talking about your
code? I used the same idea in my two-level system -- controlling the
sensed level of nutrition -- but I don't recall where my model branched
off from yours. I thought I was working with a copy of yours, but maybe
I wasn't.

Last week I gave a copy of the program to my brother, who did run it
and came back with an interesting comment I thought I'd share with CSG-
L. He thought that the graph of dNut over time looked suspiciously
like chaos.

It's not that easy to distinguish between chaos and a simple random
process. As I understand it, you have to calculate (somehow) the
dimensionality of the relationship and show that it's fractional. The
term chaos has come to be used somewhat loosely, but as I'm not an
expert I don't really know.

By the way, according to Bill's definition, this one "learns" too: the
high- level "fuel" regulating system alters the gain of the lower-level
dNut control system. (I still disagree that this is learning.)

I won't argue. I'm leaning toward reserving "learning" to mean a
structural change in the system, meaning that if we can account for
observed behavior using a system with a fixed organization, learning
isn't the proper word to use. We would use learning if that organization
were brought into being by the process, but not if it already exists.

I agree with Bill that our strategy should be to focus on terminal
performance first and worry about learning later, as appears necessary.

Good. I think we'll know when learning needs to be dealt with.

···

-------------------------------

Fortunately, how one characterizes the "VI feedback function" does not
matter if what you are going to actually do is construct a computer
model of the organism and have it interact with the actual schedule
itself and not with its representation as a proposed mathematical
formula.

I'm pleased that you agree about this. I think that a lot of the problem
in EAB is that the people in it didn't come up through the engineering
route where you learn how to model systems. There seems to be a lot of
manipulation going on that just expresses the same relationship in
another form, through equivalence transformations, and little awareness
that this doesn't change the basic relationship at all (or provide the
needed second relationship). Do you suppose that JEAB would accept a
short article on this from you? It's a major point, and a problem that's
led to a lot of wasted effort.

The problem faced by Baum and others is that their black-box philosophy
does not permit them to model the organism. Instead they are
constrained to model only the observable behavior of the organism and
the interaction of that behavior with schedule parameters. So where
are you going to put those O- rules if you have no organism to put them
in?

Exactly. Article?

In your model, there is a perceptual variable that drops steadily
between reinforcements, so that longer inter-reinforcement intervals
result in larger perceptual errors, which in turn lead to higher
response rates. Such a model would predict that the animal should
respond at maximum rates when the inter-reinforcement interval is
infinite. This might make sense if keypecking were the only source of
food, but in reality the pigeons only have to wait until the session is
over in order to receive enough grain to keep their weights stable.

By making the decay constant smaller and decreasing the reinforcement
size, you can smooth out the variations in the perceptual variable as
much as you like. The only reason I put that in was to reproduce the
scalloping and pause effects. That's really a side-issue.

The basic problem is that when the animal is on a 1:1 ratio, the
behavior rate drops nearly to zero (relatively), even though the
reinforcement rate is at maximum. If you then go to larger ratios, the
reinforcement rate falls and the behavior rate increases. That is the
behavior on the right side of the Staddon/Motherall figure.

Over on the left side we see the opposite relationship: a drop (rise) in
reinforcement rate goes with a drop (rise) in behavior rate.

A straight simple control model predicts that the relationship on the
right side should continue as we move to the left, so zero reinforcement
rate should go with maximum behavior rate. The fact that it doesn't
tells us that in the left-hand region, something is going on that
doesn't fit the simple control model.

You say "This might make sense if keypecking were the only source of
food, but in reality the pigeons only have to wait until the session is
over in order to receive enough grain to keep their weights stable."

Well, the explanation may be contained in that fact, but we have to get
it into the model. The simplest model I found that works is one in which
the rate of behavior is considered a cost and the rate of reinforcement
a benefit. We can say arbitrarily that at the point where the plot
begins to fall below the prediction of the control model (going to the
left), cost begins to exceed benefit (we can adjust one parameter so
that occurs wherever we want). Then as the cost rises above the benefit,
the difference is amplified (by a second parameter) to reduce the gain
of the control system. Staddon suggested something similar, although he
didn't follow through with the control model. This model, when the two
extra parameters are adjusted optimally, fits all the data points with
very good accuracy (a few percent).

On the other hand, the same phenomenon could be accounted for by saying
that below some rate of reinforcement, the animal simply starts spending
time elsewhere in the cage, looking for an alternative source of food.
The good fit of my proposed model may mean nothing. This is why we need
to know whether the animal is always at the bar ready to press it --
simple observation should tell us whether my model is needed.

In my view, longer inter-reinforcement intervals should lead to less
vigorous responding, not more, especially since most VI schedules tend
to generate much more responding than is necessary to obtain most or
all available reinforcers.

Look at the Staddon/Motherall data. They say otherwise. So do Skinner's
reports on shaping, where he was able to get incredibly high rates of
responding by progressively increasing the ratio and thus reducing the
reinforcement rate. I think our control model will also produce more
behavior than necessary on VI schedules. Remember that the control model
proposes that the rate of responding is driven by the difference between
the rate of reinforcement the organism wants (the reference signal) and
the amount it gets (the perceptual signal). The error signal is the
cause of the behavior. On a VI schedule, the rate of reinforcement at
high rates of behavior is going to be smaller than on an equivalent FR
schedule, leading to a large error, leading to more responding than is
necessary.

I'm about to post a general operant conditioning model that permits
selection from FR, VR, FI, and VI schedules, plotting average rates of
responding and reinforcement on one plot along with a "cumulative
record," and plotting the perceptual signal and reference signal on a
second plot. The SetParam unit is used to allow changing any parameters
while the model is running. There are some modifications to the SetParam
Unit, so I'll include it, too -- mainly to allow changing byte
parameters, which enumerated types are when there are less than 256
elements.

This model is a simple one-level model and doesn't handle all the
Staddon/Motherall data (only the five right-most points). I figure we
can match this model to real behavior over a range of schedules, and see
which parameters would be good candidates for a second level of control
to vary as a way of extending the range of application of the model.
There are too many redundant free parameters in the model as it stands
now; I'll try in the future to reduce them to just the essential ones.
------------------------------------------------------------------------
Rick Marken (941203.1330) --

That's more "natural" than my way?

Purely subjective call.

The law of effect model converges to a control model only as long as
the consequences of responses "support" this convergence.

Yes, I saw this belatedly. What I realized is that even if the results
of the previous behavior DON'T predict correctly, the organism can still
decide that it wants the less probable outcome and behave to get it.

However, let's not throw the baby out etc.. If you're at my house and
find that looking in the third cupboard from the left is accompanied by
finding a coffee-cup, the next time you look for a coffee-cup you'll
probably look where you found it the last time instead of starting the
search over. In some situations the Law of Effect is a reasonable
description of what happens. But not, of course, in all.

Despite Bill's nice attempt to rescue Bruce from the jaws of ignominy..

I guess your goal is pretty obvious. How much do _you_ enjoy ignominy?

The law of effect has no rules for "shutting off" when "learning" is
complete; if the law of effect is "true" then responses are _always_
selected by their consequences. When the environment changes so that
consequences are no longer polite enough to select the "right"
responses (the ones that result in control) then there is no more
control.

I agree that the control model embedded in the Law of Effect model
confuses the issue. If someone who had never seen the control-system
model had tried to apply the Law of Effect, the outcome might have been
different.

I think you're overlooking the fact that the Law of Effect model was
used to vary a parameter of the control system, not to accomplish the
control all by itself. That's why Bruce referred to it as a learning
model. The basic organization, where dNut influenced the delay between
tumbles, was already given. Only the direction and amount of the effect
was altered by the logical model. I described several other models that
would, in principle, have the same effect, starting with no probability
of tumbling. So I don't believe that the LOE model is unique in being
able to cause this sort of "learning." However, it did work under the
stated conditions, so what's the problem? Accepting that doesn't require
accepting the LOE as a basic principle, does it?

I think that the valid points you want to make can be made much more
clearly and easily in an operant-conditioning model where we don't have
that pesky random output to confuse matters, or even more simply in a
simple control experiment. One model is going to fit the data better
than others and do it over a wider range. I think it will be clear which
model that is by the time we're done.

It seems to me that it would be impossible to get the right _feel_ for
PCT theory and research if one imagined that the behavior of a living
control system could be legitimately conceived of as selected BY it
consequences.

I agree. But one has to be familiar with the control model and see it in
operation before this can be understood. For anyone to agree with your
conclusion before understanding why it is right would be a mistake,
wouldn't it? Should we tell people to stop trying to understand _why_
and just take our word for it? Learn their catechism? Memorize the
answers in the back of the book? Feed back to us what we tell them, to
get an "A"?

To get here, everybody must start somewhere else. We still have a lot to
learn about how to serve as guides on the journey.
------------------------------------------------------------------------
Best to all,

Bill P.