[From Bill Powers (941119.0700 MST)]
Rick Marken (941118.1345) --
So I guess I'll either have to 1) come up with a better demonstration
that Bruce's reinforcement doesn't control or 2) show that Bruce's
model is really a control model or 3) give up and admit that
reinforcement and control theory are just two alternative models of the
same thing -- purposeful behavior.
3) isn't the only final alternative. Actually, if an observable
consequence of behavior _could_ act on the organism so as to increase
(or decrease) the probability of the behavior that creates that
consequence, Bruce's model would describe that effect correctly (in one
of the possible ways). Not only that, the behavior would become more (or
less) frequent and we would see the results that we see.
Bruce's model was very clever. It took the basic control model and split
it into two parts, one for positive dNut and another for negative dNut.
Then it modeled a mechanism which linked the change in dNut across a
tumble to an increase in the delay of the part associated with positive
dNut, or a decrease in the delay of the other part. Since the aspect of
the behavior affected by the reinforcement was a probability, it had
natural limits at 0 or 1, so runaway effects of the positive feedback
(more reinforcement, more behavior, more reinforcement...) were
prevented. Each branch of the loop approached its limit and stayed
there, with one probability at 1 and the other at 0. In this state, the
whole system was equivalent to a single control system with high gain
and a reference level of zero: "if dNut < 0 then DoTumble." The "high
gain" aspect comes from the fact that dNut only has to be
infinitesimally less than the reference value of 0 to produce a tumble.
So this was a successful model of some organism. Some questions,
however, remain as to whether it is actually a reinforcement model and,
regardless of that, whether it is a model of a real organism.
The word reinforcement has been around for so long that it has become a
thing in its own right and we tend not to think of what it means. A
reinforcer is an observable physical something that is supposed to
reinforce a given behavior: that is, it has an effect on the organism
that a non-reinforcer like a piece of gravel would not have; to increase
the probability that there will be a given response to a given stimulus.
But what property does a piece of matter or an external situation have
to have in order to be able to create this effect? Reinforcement theory
is silent on this question. In fact, reinforcement is a dormitive
principle: a sleeping powder makes you sleepy because it contains a
Dormitive Principle; a piece of food increases the probability of
behavior because it contains a Reinforcing Principle. A piece of gravel
does not change behavior because it lacks the Reinforcing Principle.
In Bruce's model, the change in dNut across a tumble is said to have
reinforcing properties. However, no such properties have been given to
that change. Instead, we find inside the organism a complex mechanism
which takes concentration as an input, computes the rate of change of
concentration, computes the rate of change of that, and uses the outcome
of the calculation to adjust a parameter in a probability calculation
upward or downward. Then another calculation continually generates a
random number and compares it to the probability parameter, and if the
result is negative, produces a random tumble, still another mechanism.
So the change in behavior we see is not due to the objective changes in
rates of change of concentration, but to a series of computational
mechanisms proposed as part of the model of the organism. If the
calculations attributed to the organism were different, the outcome
would be different, even without any change in the physical properties
of the environment. Yet if the systematic behavioral changes ceased, it
would be said, under reinforcement theory, that the changes in
concentration had lost their reinforcing properties.
When a Principle is observed to lose its effect, the natural solution is
to propose another Principle that has the opposite effect and to say
that it has come to predominate. So if an object ceases to have
reinforcing effects, it must be because the same object contains a
Satiation Principle that comes into play only at high levels of
reinforcement. Or it might be that high levels of reinforcement contain
an Aversive Principle, equivalent to a Punishing Principle, leading to
an decrease in the probability of the same behavior that is increased
when the Reinforcing Principle predominates.
Bruce's model is not an example of reinforcement theory. The changes in
behavior that are seen are not due to reinforcing properties of dNut or
its changes, but to mechanisms inside the modelled organism without
which neither dNut nor any of its derivatives would have any effect at
all.
···
----------------------------
Which leads us to the other question: is Bruce's model a plausible model
of a real organism? There is no question that it works, but is there a
simpler way, a simpler mechanism that we can propose for the insides of
the organism, that will create the same effect? The answer is clearly
yes. We can do completely without the calculation of the change in dNut
across a tumble. We need no probability calculations. Using about 1/3 of
the program code, we can not only explain the behavior originally seen,
but with the addition of four or five more lines of code, show how E.
coli could also seek a particular level of concentration, neither
lethally low nor lethally high. So on grounds of parsimony, we can say
that a pure PCT model is preferable to one in which far more complex
calculations are attributed to the organism.
----------------------------
In the E. coli model there is a basic random process, the tumbles, which
make more complex aspects of behavior hard to understand. When we turn
to the operant conditioning models as I hope we can now do, we will find
a simpler situation in which we can ask even more interesting questions
about reinforcement theory. We will be able to show clearly that the
rate of reinforcement is not an independent variable. We will be able to
show that disturbances tending to change the rate of reinforcement are
resisted by changes in the action of the organism, but that changes in
the action caused by disturbances are not resisted by changes in the
reinforcement rate. In this way we will be able to show that
reinforcement rate is under the control of behavior, and that behavior
is not under the control of reinforcement rate. This will be a less
philosophical and more direct refutation of the basic concept of
reinforcement. And thanks to Bruce, we will be able to demonstrate each
step of this proof by experiments with real animals.
-----------------------------------------------------------------------
Best,
Bill P.