[From Bill Powers (970320.0445 MST)]
Bruce Abbott (970319.1920 EST)--
First, in reinforcement theory, consequences don't "select responses" any
more than, in Darwinian theory, consequences "select genes." Rather,
those responses that are followed by certain sorts of consequences tend to
occur more often, under the conditions in which they had been followed by
those consequences in the past, just as the characteristics of animals
that reproduce better tend to occur more often in the population in
succeeding generations.
If you can't leave this alone, neither can I.
The assumption under reinforcement theory is that the _responses_ are what
become more frequent. But we know from PCT that in a normal world, if you
repeat the same responses you will not generate the same consequences.
Usually, in order to generate a specific consequence, you must _vary_ your
"responses" (actions), sometimes only by a little but sometimes by a great
amount. And organisms do this all the time. We call the process "controlling."
The PCT view would be that it is the _consequence_ that becomes more
frequent, once it has been seen to occur, and if it is the sort of
consequence for which the organism has a nonzero reference level. In
learning to recreate the desired consequence, the organism varies its
behavior until the consequence occurs again. During this process of learning
to control the consequence, the organism approaches the critical act from
many starting points and orientations, so it is learning how to vary its
actions to compensate for the disturbances introduced by these changes in
initial conditions. It thus learns not just one but several control
processes which are needed as part of the overall process of controlling the
desired consequence. It learns how to act to oppose errors over a range
around zero. No particular action becomes more likely; if it did, that would
prevent compensating for some of the variations in initial conditions.
This picture is obscured when the organism is in a situation where, in fact,
only one action will produce the desired consequence (a rarity in the
natural world, if not in the laboratory). Then, of course, an increase in
the frequency of occurrance of the consequence must be accompanied by an
increase in the frequency of occurrance of the only action that can produce
that consequence. It then becomes impossible to tell, experimentally,
whether it is the frequency of the consequence or the frequency of the
action that is being increased. The two frequencies necessarily covary.
In order to tell which is the proper description, we must introduce
appropriate disturbances. If we observe that a _different_ action is used to
produce the _same_ consequence, as required by the disturbance, we have then
proven that it is the consequence, not the action, that is coming under
control, and that it is the action, not the consequence, that is the means
of control.
Suppose a person is just learning how to use a mouse to bring a cursor to a
target. If each trial begins with the cursor and mouse in a specific
position, and if there is no disturbance, then there is only one motion of
the mouse that will (with minimum effort) accomplish the intended result of
cursor on target. We will observe, after enough practice, a particular
motion of the mouse that always brings the cursor to the target. So it could
be said that the movement of the cursor to the target is the reinforcing
consequence that causes, over many trials, the generation of the mouse
movement that is observed. We could see the condition "cursor on target" or
"cursor moving to target" as a reinforcing state of affairs, and the mouse
movement as being conditioned by that consequence, the discriminative
stimulus being the initial position of the cursor relative to the target. In
fact, by using different starting conditions, we could show that the mouse
movement becomes just what is required to bring the cursor to the target
from each starting position, so the behavior is brought under the control of
the different discriminative stimuli by the reinforcing consequence of mouse
movement.
But now let's introduce disturbances. At the start of each trial, an
invisible disturbance begins to affect the cursor position along with the
effects of mouse movements. If the disturbance begins pushing the cursor
away from the target, the mouse movement will become greater. If the
disturbance pushes the cursor toward the target, the mouse movement will
become less -- in fact, it could reverse, if the disturbance becomes large
enough by itself to push the cursor _beyond_ the target. So now we have the
same "discriminative stimuli" (initial conditions), and the same
"reinforcing consequence" (cursor moving to target position), but different
actions on every trial: the mouse moves by different amounts and even in
different directions. Indeed, by choosing the disturbances carefully, we
could make the mouse movements appear to be randomly distributed. Yet on
every trial, the cursor would move to the target position.
It seems to me that the extra information we get from introducing the
disturbance clearly distinguishes between reinforcement theory and control
theory. The reinforcement explanation is not simply another possible point
of view: it is wrong -- at least for the tracking experiment. To see if it
is wrong for the operant-conditioning cage, we would have to introduce
appropriate disturbances, to see if the reinforcing event continues to occur
even though different actions are needed to bring it about under the same
discriminative stimuli. This, of course, would require a redesign of the
experiments, because their normal design is such that the reinforcer can be
produced only by specific actions; as a result, the frequency of occurrance
of reinforcements and actions covaries. Without disturbances acting, the
direction of control is ambiguous, and indeed, which explanation you use is
a matter of preference, neither being justifiable.
Back to E. coli.
Second, how movement away from the target can be considered reinforcement
in the demo is beyond my understanding. Responses that decrease movement
away from the target or increase movement toward the target would be
expected to occur more often over time (reinforcement). This means that
responses made while the bug is moving away from the target will tend to
increase in frequency, given that such responses are more often followed
by an improvement than a worsening (as is the case in the demo). The
opposite would be expected for responses made while the bug is approaching
the target (i.e., the frequency of such responses would tend to decrease
over time, a phenomenon known as punishment).
If you will recall our explorations of your model for E. coli (although you
disclaimed any attempt to model the actual bacterium, as well you might
considering the logical calculations it carried out), the actual explanation
did not turn out to be the one you offer above. It was necessary for the
model to discriminate four logical conditions, and only two of them worked
in the right direction. These were both _negative_ effects (or perhaps more
correctly, punishments). The positive effects actually ended up working in
the wrong direction, but they were outweighed by the negative effects. What
was actually required was a _decrease_ in the probability of a tumble when
the controlled concentration was _increasing_, and an _increase_ in
probability of a tumble when the concentration was _decreasing_. That, of
course, is simply the control-system model we proposed, which works without
any probability calculations.
However, two of the conditions in your model required an _increase_ in
probability of a tumble when concentration was _increasing_ and a decrease
when it was decreasing, just the opposite of what is needed. What saved it
was the negative effects: if the effect of a tumble was to make matters
worse, the probablity of a tumble under those conditions was _decreased_. I
suppose these were not really "negative reinforcers" but actual punishments.
Those effects turned out to be somewhat stronger than the incorrect ones, so
the model did work.
Note that the reinforcement model is a learning model; behavior that is
initially random with respect to the target becomes keyed to the initial
direction of movement, until (in the limit) responses occur when the bug
is heading in the wrong direction and not when it is heading in the right
direction. The control model in this demo doesn't have to learn anything,
because it is set up "knowing" what to do to correct any error, right out
of the box.
I disagree. If you changed the attractant into a repellant, you would have
to go into your model and manually redefine which alternative is the good
one and which is the bad one. Your model was set up to prefer going up a
gradient to going down it, no matter what the substance was. The "learning"
you see as a "change in probability of a tumble" was simply a roundabout way
of making the delay to the next tumble longer if the direction of travel was
up the gradient, etc. The control system model works this way, too, except
that the error signal simply alters the delay directly, instead of doing it
by way of a random number generator. And your model, because it worked
partly in the wrong direction, went up the gradient more slowly than the
control model did.
Now before anyone thinks that I am defending reinforcement theory, think
again. From one perspective, "reinforcement" is only a term that refers
to what is observed as a control system gets organized and tuned.
But your model did not reorganize itself. It was set up so that, taking into
account both the incorrect and the correct adjustments it made, it delayed
the tumbles when moving up the gradient and generated them sooner when
moving down the gradient. This result was obtained in a very complicated
way, but in a fixed way. No matter what happened, it moved up the gradient,
if slowly. You might go through the program and change every reference to
"Nut" (nutrient) to "Rep" (repellant), but when you compiled and ran the
program it would still move up the gradient. To get it to move down the
gradient of a repellant, you would have to rewire the logic, and your model
had no way of doing that (of course the control model couldn't do it,
either). YOU had to tell your model that a positive change in dNut was good.
And YOU had to provide the logic that made it move up the gradient.
What you are seeing as reorganization in your model was simply an inevitable
increase in the loop gain, as the relative probabilities changed in the only
direction they could change. No matter where you started, you ended up with
the same probabilities: this is hardly "learning." It would be learning only
if some other outcome was possible.
Best,
Bill P.