In Search of Learning

[From Bruce Abbott (941127.1715 EST)]

I've been in electronic solitary confinement since late Wednesday night, when
I received a post from Bill Powers entitled "Retraction," a brief message
promising more in the morning. Our internet connection evidently went dead
shortly thereafter, not to be restored until past noon today (Sunday). On
downloading the log file, wow. It will take me a couple of days just to make
hard copies of all the posts and read them. A quick scan indicates that the
Enterprise shields are once again up. I'd hoped that by now I'd be beaming
aboard.

Bill Powers (941123.1415 MST)

I don't know what you will make of this, but in fact the logical path
you describe just above is NOT the effective path-- it works the wrong
way, with a probability of 0.37. The effective path (probability 0.63)
is

     If the previous action made S(new) WORSE than S(old) then

     If s(old) was favorable, DECREASE the probability of the action
     that takes place when s(new) is favorable.

You don't know what I will make of this? This is what I've been trying to
TELL you! Although the results of a tumble under this discriminative stimulus
are sometimes favorable (more nutrients), more often than not, a tumble makes
things worse. Thus, on average, tumbling is punished and so occurs less and
less often. Let's see whether this result agrees with intuition:

Things are going along O.K. for you, but you try something and things get even
better. Are you encouraged to try it again or discouraged? [encouraged] Is
this consistent or inconsistent with the supposed effect of reinforcement on a
response? [consistent]

Things are going along O.K. for you, but you try the same something and things
get worse. Are you encouraged to try it again or discouraged? [discouraged]
Is this consistent or inconsistent with the supposed effect of punishment on a
response? [consistent]

Now let's say that, in the long run under favorable conditions, things get
better with p = .25 and things get worse with q = .75. In the long run, are
you going to be making this response more and more, or less and less, under
that favorable condition? [You are more often discouraged than encouraged, so
you make the response less and less often.]

[By the way, the conditional probabilities are .25 and .75, not .37 and .63.]

Exactly parallel logic leads to a net increase in the probability of a tumble
when things are unfavorable:

Things are going badly for you, but you try something and things get even
worse. Are you encouraged to try it again or discouraged? [discouraged] Is
this consistent or inconsistent with the supposed effect of punishment on a
response? [consistent]

Things are going badly for you, but you try the same something and things get
better. Are you encouraged to try it again or discouraged? [encouraged] Is
this consistent or inconsistent with the supposed effect of reinforcement on a
response? [consistent]

You are encouraged with p = .75 and discouraged with q = .25, so your response
becomes more and more likely.

So the model is entirely consistent with the law of effect as stated by
Thorndike over 80 years ago, as you now agree:

So in this situation we have to admit that the Law of Effect model as
you represent it does provide an explanation of the behavior of E. coli

But this result capitalizes on a "peculiar situation," you say, as if somehow
that fact invalidates the law of effect. So what? E. coli learns what it
learns because that is what the environment teaches. Is this "peculiar?"

Your next step is to propose a control model that supposedly learns. Here we
get into grave difficulty, because neither of us have taken the trouble to
carefully define what we mean by "learning." Without a clear definition on
which we both agree, we cannot objectively determine whether we have a model
that learns or not.

But learning is a tricky beast to define, as evidenced by the large number of
proposals made over the past century. For the current purpose I am going to
opt for a simple one: learning is a change in behavior as a result of past
experience.

I say "past" experience in order to differentiate the effect of previous
experience from that of present experience. To learn, an organism must have a
past whose effects modify the behavior system in some consistent way. Given a
different history of experience, the organism should behave differently.

I maintain that ECOLI4a meets this definition. If the "peculiar" environment
were modified in such a way that tumbling during favorable conditions led to
even better conditions more often than not (opposite to what happens now), and
tumbling during unfavorable conditions led to even worse conditions more often
than not (also opposite to what happens now), this model would quickly learn
to tumble during favorable conditions and not to tumble during unfavorable
conditions. ECOLI4a's behavior in the presence of favorable and unfavorable
conditions adapts to the net contingencies as a function of EXPERIENCE with
the consequences of its own behavior.

Let us now examine your model to see whether it, too, is capable of such
adaptive changes in behavior as a function of experience. Here is your model:

                          Ref
                       - +|
         Nut -->dNut --->Comp --> Gain --> Delay--- TUMBLE
          > +| |-
          >--->Effect of Nut good-> | increment or
           --->Effect of Nut bad ->-- decrement gain

Here the delay depends on the error signal which depends on the current
value of dNut; if the gain is positive, dNut will generally be kept
positive. In fact this model uses only present-time values of all
variables.

We start out with the gain at zero, so e. coli's lower-level control system
has no effect on the rate of tumble, which occurs at some base rate. Let us
assume that, in this instance, e. coli is advancing up the nutrient gradient.
The change in nutrient, dNut, is positive, but this is having no effect on
behavior because the gain is currently zero. The incoming nutrients are
apparently improving (reducing error in) some higher-level perception. What
should our ignorant little e. coli do? Should it suppress the rate of
tumbling or increase it? [Remember, it does not know what effect a tumble
will have on the rate of nutrient intake.] But what does your model say? It
says that, so long as the "effect of Nut good" state persists, make the gain
more positive, i.e., suppress tumbling. Now, I ask you--how did e. coli know
that this is the solution to its dilemma? The answer is that you PROGRAMMED
that response into it. In other words, it was BORN with the knowledge, it did
not acquire it via experience with the consequences of its behavior. It did
not learn.

This truth is most easily seen if you imagine a new environment, in which
tumbling during favorable conditions made things better more often than not,
and tumbling during unfavorable conditions made things worse more often than
not. Your model would simply persist in its old behavior until it starved to
death (because increased dNut is still "good" and thus produces positive gain,
which decreases the tumble rate.) As I described earlier, ECOLI4a would adapt
nicely to this environment, changing its behavior to suit the new
contingencies.

There can be no learning from the consequences of one's behavior unless one
(a) tries the behavior, (b) notes the consequence, and (c) uses this
information to alter future behavior. If your model uses only current values
of perceptual variables to guide its current behavior, it is not learning.
Thus I do not agree that:

But it is also true that this behavior can be explained
at least equally well without assuming that reinforcement is taking
place or that previous values of dNut figure into the behavior at all.

[Rick: I've scanned your two posts in which you attempt to "rig" the
environment so that ECOLI4a does not learn the appropriate behavior. When you
succeed, it is because you have broken the contingency between behavior and
its consequences. What you are saying is that when there is nothing useful to
learn from the consequences of one's behavior, one does not learn anything
useful from one's behavior. This is precisely the behavior you would expect
from a learning e. coli!]

Clearly the control model assumes less and is simpler. We have two
models to explain the same phenomenon, and neither one can be
transformed into the other: they are not equivalent. To accept the
control model is therefore to say that the behavior of E. coli is _not_
governed by reward and punishment, and that it does NOT depend on
previous values of dNut or on the change of dNut across a tumble.

I, too, believe in parsimony as one test of a theory's merit. But to say
that, in the control model, the behavior of a LEARNING e. coli does not depend
on experience (previous values of dNut, change in dNut across a tumble)
appears to me to say that learning does not depend on past experience, a
contradiction.

When two models predict the same behavior equally well but are not
equivalent to each other, they are making different claims about the
internal organization of the behaving system. I am of the school that
considers such a situation to be unacceptable. The only thing to do
before either model is actually used is to find some way of testing the
different claims to see which are verifiable.

We must have gone to the same school. Before we can determine whether your--
or my--model learns, we must first agree on what constitutes learning. Then
and only then can we agree whether or not a given model learns.

So where are we now? Do we agree on a provisional definition of learning? If
so, can you show me where learning is in your e. coli model? I don't see it.

On rereading your post, I get the feeling that we are, once again, arguing
different issues. In part (particularly near the end of your post), you are
arguing that PCT provides a model of the behaving system superior to that of
to traditional reinforcement theory, to which I say, OF COURSE! No QUESTION!
My argument concerns how effective control systems come into being. I am
proposing a fast, efficient method which, though couched in the antiquated
terminology of an outdated 1911 reflex model, uses the principles of variation
and selection to construct appropriate behavioral responses to perceptual
errors.

Regards,

Bruce

P.S. Your Next Generation parody (temporarily lost to me because of internet
problems but now retrieved) had me in rolling on the floor for about an hour--
and it gave me an idea. I've been offering my dog Muffie a treat for jumping
up and grabbing a stick out of my hand. I started by holding the stick at
chest level and gradually moved it higher. At one point I had to get a
ladder. When she reached the stick from that height I moved to the roof of
the house. Then I began to withhold the treat until she stayed at stick-
height for at least one second, and gradually increased the time requirement.
She's up to 12 seconds now. With a little more work, I think she'll be able
to duplicate the behavior of your blue fuzzies--although getting her to light-
speed may take a bit more doing.

Greet us (;->