Marken JAVA Demos

[From Bruce Abbott (970319.1920 EST)]

I've taken a look at Rick's JAVA demos and have a couple of comments to
make. First, Rick, you might want to refer to the device used to control
the CV by name (e.g., mouse); "handle" is a term Bill used long ago when the
device literally was a handle, and for some reason this term has persisted.
I doubt that many readers will know what you mean by that term.

My second, and more extensive comment has to do with the "Selection by
consequences" demo. Below I quote Rick's description of "reinforcement":

The consequences that are supposed to be selecting responses are called
reinforcements. A reinforcement is something (like money) that increases
the probability of the response that produces it. In this study reinforcement
is movement of the bug relative to the target: movement away from the target
is strongest reinforcement because it results in the largest increase in the
probability of the response (bar press) that produces it and movement toward
the target is the weakest reinforcement because it results in no increase in
the probability of the response that produced it.

First, in reinforcement theory, consequences don't "select responses" any
more than, in Darwinian theory, consequences "select genes." Rather, those
responses that are followed by certain sorts of consequences tend to occur
more often, under the conditions in which they had been followed by those
consequences in the past, just as the characteristics of animals that
reproduce better tend to occur more often in the population in succeeding
generations.

Second, how movement away from the target can be considered reinforcement in
the demo is beyond my understanding. Responses that decrease movement away
from the target or increase movement toward the target would be expected to
occur more often over time (reinforcement). This means that responses made
while the bug is moving away from the target will tend to increase in
frequency, given that such responses are more often followed by an
improvement than a worsening (as is the case in the demo). The opposite
would be expected for responses made while the bug is approaching the target
(i.e., the frequency of such responses would tend to decrease over time, a
phenomenon known as punishment).

Rick and Bill and I went over this rather thoroughly a couple of years ago,
and I was able to prove via computer simulation that a _correct_
reinforcement-based model converged on the target just as the control-system
model did. To put it bluntly, Rick's description of the reinforcement model
and his "demonstration" of its behavior in this situation is incorrect, and
he knows it.

Note that the reinforcement model is a learning model; behavior that is
initially random with respect to the target becomes keyed to the initial
direction of movement, until (in the limit) responses occur when the bug is
heading in the wrong direction and not when it is heading in the right
direction. The control model in this demo doesn't have to learn anything,
because it is set up "knowing" what to do to correct any error, right out of
the box. When you try the task yourself, you, of course, must learn what
"works" and what doesn't before you gain control and get the bug moving
generally toward the target.

Now before anyone thinks that I am defending reinforcement theory, think
again. From one perspective, "reinforcement" is only a term that refers to
what is observed as a control system gets organized and tuned. As such it
not an explanation of those changes: those actions that tend to reduce
error, or reduce it more effectively, get "selected" over those that
increase or fail to decrease error, or reduce error less effectively. The
control system thereby develops actions that tend to counter disturbances,
and an appropriate level of gain.

I don't think that PCT is so weak that it can demonstrate its advantages
over reinforcement theory only by misrepresenting the latter. Rick, how
about giving that demo the pitch? Pretty please?

Regards,

Bruce

[From Bill Powers (970320.0445 MST)]

Bruce Abbott (970319.1920 EST)--

First, in reinforcement theory, consequences don't "select responses" any
more than, in Darwinian theory, consequences "select genes." Rather,
those responses that are followed by certain sorts of consequences tend to
occur more often, under the conditions in which they had been followed by
those consequences in the past, just as the characteristics of animals
that reproduce better tend to occur more often in the population in
succeeding generations.

If you can't leave this alone, neither can I.

The assumption under reinforcement theory is that the _responses_ are what
become more frequent. But we know from PCT that in a normal world, if you
repeat the same responses you will not generate the same consequences.
Usually, in order to generate a specific consequence, you must _vary_ your
"responses" (actions), sometimes only by a little but sometimes by a great
amount. And organisms do this all the time. We call the process "controlling."

The PCT view would be that it is the _consequence_ that becomes more
frequent, once it has been seen to occur, and if it is the sort of
consequence for which the organism has a nonzero reference level. In
learning to recreate the desired consequence, the organism varies its
behavior until the consequence occurs again. During this process of learning
to control the consequence, the organism approaches the critical act from
many starting points and orientations, so it is learning how to vary its
actions to compensate for the disturbances introduced by these changes in
initial conditions. It thus learns not just one but several control
processes which are needed as part of the overall process of controlling the
desired consequence. It learns how to act to oppose errors over a range
around zero. No particular action becomes more likely; if it did, that would
prevent compensating for some of the variations in initial conditions.

This picture is obscured when the organism is in a situation where, in fact,
only one action will produce the desired consequence (a rarity in the
natural world, if not in the laboratory). Then, of course, an increase in
the frequency of occurrance of the consequence must be accompanied by an
increase in the frequency of occurrance of the only action that can produce
that consequence. It then becomes impossible to tell, experimentally,
whether it is the frequency of the consequence or the frequency of the
action that is being increased. The two frequencies necessarily covary.

In order to tell which is the proper description, we must introduce
appropriate disturbances. If we observe that a _different_ action is used to
produce the _same_ consequence, as required by the disturbance, we have then
proven that it is the consequence, not the action, that is coming under
control, and that it is the action, not the consequence, that is the means
of control.

Suppose a person is just learning how to use a mouse to bring a cursor to a
target. If each trial begins with the cursor and mouse in a specific
position, and if there is no disturbance, then there is only one motion of
the mouse that will (with minimum effort) accomplish the intended result of
cursor on target. We will observe, after enough practice, a particular
motion of the mouse that always brings the cursor to the target. So it could
be said that the movement of the cursor to the target is the reinforcing
consequence that causes, over many trials, the generation of the mouse
movement that is observed. We could see the condition "cursor on target" or
"cursor moving to target" as a reinforcing state of affairs, and the mouse
movement as being conditioned by that consequence, the discriminative
stimulus being the initial position of the cursor relative to the target. In
fact, by using different starting conditions, we could show that the mouse
movement becomes just what is required to bring the cursor to the target
from each starting position, so the behavior is brought under the control of
the different discriminative stimuli by the reinforcing consequence of mouse
movement.

But now let's introduce disturbances. At the start of each trial, an
invisible disturbance begins to affect the cursor position along with the
effects of mouse movements. If the disturbance begins pushing the cursor
away from the target, the mouse movement will become greater. If the
disturbance pushes the cursor toward the target, the mouse movement will
become less -- in fact, it could reverse, if the disturbance becomes large
enough by itself to push the cursor _beyond_ the target. So now we have the
same "discriminative stimuli" (initial conditions), and the same
"reinforcing consequence" (cursor moving to target position), but different
actions on every trial: the mouse moves by different amounts and even in
different directions. Indeed, by choosing the disturbances carefully, we
could make the mouse movements appear to be randomly distributed. Yet on
every trial, the cursor would move to the target position.

It seems to me that the extra information we get from introducing the
disturbance clearly distinguishes between reinforcement theory and control
theory. The reinforcement explanation is not simply another possible point
of view: it is wrong -- at least for the tracking experiment. To see if it
is wrong for the operant-conditioning cage, we would have to introduce
appropriate disturbances, to see if the reinforcing event continues to occur
even though different actions are needed to bring it about under the same
discriminative stimuli. This, of course, would require a redesign of the
experiments, because their normal design is such that the reinforcer can be
produced only by specific actions; as a result, the frequency of occurrance
of reinforcements and actions covaries. Without disturbances acting, the
direction of control is ambiguous, and indeed, which explanation you use is
a matter of preference, neither being justifiable.

Back to E. coli.

Second, how movement away from the target can be considered reinforcement
in the demo is beyond my understanding. Responses that decrease movement
away from the target or increase movement toward the target would be
expected to occur more often over time (reinforcement). This means that
responses made while the bug is moving away from the target will tend to
increase in frequency, given that such responses are more often followed
by an improvement than a worsening (as is the case in the demo). The
opposite would be expected for responses made while the bug is approaching
the target (i.e., the frequency of such responses would tend to decrease
over time, a phenomenon known as punishment).

If you will recall our explorations of your model for E. coli (although you
disclaimed any attempt to model the actual bacterium, as well you might
considering the logical calculations it carried out), the actual explanation
did not turn out to be the one you offer above. It was necessary for the
model to discriminate four logical conditions, and only two of them worked
in the right direction. These were both _negative_ effects (or perhaps more
correctly, punishments). The positive effects actually ended up working in
the wrong direction, but they were outweighed by the negative effects. What
was actually required was a _decrease_ in the probability of a tumble when
the controlled concentration was _increasing_, and an _increase_ in
probability of a tumble when the concentration was _decreasing_. That, of
course, is simply the control-system model we proposed, which works without
any probability calculations.

However, two of the conditions in your model required an _increase_ in
probability of a tumble when concentration was _increasing_ and a decrease
when it was decreasing, just the opposite of what is needed. What saved it
was the negative effects: if the effect of a tumble was to make matters
worse, the probablity of a tumble under those conditions was _decreased_. I
suppose these were not really "negative reinforcers" but actual punishments.
Those effects turned out to be somewhat stronger than the incorrect ones, so
the model did work.

Note that the reinforcement model is a learning model; behavior that is
initially random with respect to the target becomes keyed to the initial
direction of movement, until (in the limit) responses occur when the bug
is heading in the wrong direction and not when it is heading in the right
direction. The control model in this demo doesn't have to learn anything,
because it is set up "knowing" what to do to correct any error, right out
of the box.

I disagree. If you changed the attractant into a repellant, you would have
to go into your model and manually redefine which alternative is the good
one and which is the bad one. Your model was set up to prefer going up a
gradient to going down it, no matter what the substance was. The "learning"
you see as a "change in probability of a tumble" was simply a roundabout way
of making the delay to the next tumble longer if the direction of travel was
up the gradient, etc. The control system model works this way, too, except
that the error signal simply alters the delay directly, instead of doing it
by way of a random number generator. And your model, because it worked
partly in the wrong direction, went up the gradient more slowly than the
control model did.

Now before anyone thinks that I am defending reinforcement theory, think
again. From one perspective, "reinforcement" is only a term that refers
to what is observed as a control system gets organized and tuned.

But your model did not reorganize itself. It was set up so that, taking into
account both the incorrect and the correct adjustments it made, it delayed
the tumbles when moving up the gradient and generated them sooner when
moving down the gradient. This result was obtained in a very complicated
way, but in a fixed way. No matter what happened, it moved up the gradient,
if slowly. You might go through the program and change every reference to
"Nut" (nutrient) to "Rep" (repellant), but when you compiled and ran the
program it would still move up the gradient. To get it to move down the
gradient of a repellant, you would have to rewire the logic, and your model
had no way of doing that (of course the control model couldn't do it,
either). YOU had to tell your model that a positive change in dNut was good.
And YOU had to provide the logic that made it move up the gradient.

What you are seeing as reorganization in your model was simply an inevitable
increase in the loop gain, as the relative probabilities changed in the only
direction they could change. No matter where you started, you ended up with
the same probabilities: this is hardly "learning." It would be learning only
if some other outcome was possible.

Best,

Bill P.

[From Rick Marken (970320.0730 PST)]

Bruce Abbott (970319.1920 EST) --

I've taken a look at Rick's JAVA demos and have a couple of
comments to make.

Hmmm. Why am I thinking this is not going to be unrestrained praise;-)

you might want to refer to the device used to control the CV by
name (e.g., mouse); "handle" is a term Bill used long ago...I doubt
that many readers will know what you mean by that term.

Now there's an important point! Actually, I generally tried to refer to the
mouseas a "mouse". I will be happy to purge the word "handle" from my
write-ups wherever it occurs. Could you please tell me where you found it?

My second, and more extensive comment has to do with the
"Selection by consequences" demo.

Now there's an interesting mistake. It's the "Selection OF Consequences"
demo, Bruce. I guess it's as tough for an old die-hard behaviorist to say
"Selection of consequences" as it is for an old die-hard handle turner to
say "mouse":wink:

First, in reinforcement theory, consequences don't "select
responses"

Gee. I thought Skinner's "Selection by consequences" article in _Science_
(the one that suggested the title for my demo -- and paper) was about how
consequences _select_ responses? I guess Skinner didn't understand
reinforcement theory, eh? Sure wish you had explained it to him, Bruce.

Second, how movement away from the target can be considered
reinforcement in the demo is beyond my understanding.

Well, all my teachers in graduate school (know-nothings like David Premack)
told me that a reinforcing consequence is one that increases the probability
of the response that produces it. In the demo, if dot movement away from the
target is the consequence of a key press, the probability of another press
is higher than if the consequence of a key press had been dot movement toward
the target. In fact, by making "dot movement away" contingent on every
response I can dramatically increase the probability of responding -- just
what one would want in a reinforcer. I can "strengthen" the pressing
response by having the dot move away from the target after every press.

Rick and Bill and I went over this rather thoroughly a couple of
years ago, and I was able to prove via computer simulation that
a _correct_ reinforcement-based model converged on the target
just as the control-system model did.

As I recall, we all agreed that your "_correct_ reinforcement-based
model" was just an arcane version of a control-system model.

To put it bluntly, Rick's description of the reinforcement model
and his "demonstration" of its behavior in this situation is
incorrect

But other than that, Mr. Skinner, how'd you like the demo;-)

Now before anyone thinks that I am defending reinforcement theory,

Well, if you were "prosecuting" (rather than defending) reinforcement
theory then this was the worst exhibition of prosecutorial skills
I've seen since the OJ trial;-)

I don't think that PCT is so weak that it can demonstrate its
advantages over reinforcement theory only by misrepresenting
the latter.

Well, we certainly agree on this!

Rick, how about giving that demo the pitch? Pretty please?

Thank you Bruce! I'm honored. This is the best response to my demos yet.
I consider a demo a success when I see a conventional psychologist
walking away from it muttering things like "straw man" or "he doesn't
understand the theory" or "so what".

But I am kind of disappointed that you don't want me to pitch any of the
other demos. How about the "S-R vs Control" demo. Surely you take exception
to my claim that the input in this task is _not_ the cause of the outputs.
I mean we're talking about the foundations of scientific psychology here.
And how about the "Behavioral Illusion" demo. Surely you take exception to my
claim that this demo shows the futility of trying to understand behavior
using conventional IV-DV methods. I mean, this is about textbook contracts,
Bruce. Big _bucks_ are at stake;-)

Best

Rick