E. Coli Revisionism

[From Bruce Abbott (950611.1635 EST)]

Bill Powers (950609.1600 MDT) --

Rick tells me, Bruce, although the post hasn't reached here, that you
have referred back to the E. coli programs to show a successful
application of the principle of reinforcement. I rather suspected you
would, and hoped you wouldn't.

I think you have serious misunderstandings about why I wrote those programs,
and what I believe I was able, ultimately, to demonstrate. I hope I can
clear these misunderstandings up.

Be forewarned that I intend to be as of critical of your reasoning as you
have been of mine. I'm not doing this to be nasty, but since you have felt
compelled to "call 'em as you see 'em," I have felt at liberty to do the
same. Fair is fair.

I. The Challenge

The challenge was to build a model, consistent with reinforcement theory,
that would behave "properly" under special conditions designed to thwart a
reinforcement interpretation. Those conditions were as follows:

    1. The only behavior permitted was straight-line motion or "tumbling."

    2. The outcome of a tumble was to sample a new direction for
        straight-line motion. The new direction following a tumble was to be
        chosen strictly at random.

    3. The only sensory inputs available were nutrient density and its
        derivatives.

    4. The test environment would consist of a two-dimensional field and,
located in that field, a nutrient source, with nutrient concentration
        declining with distance in every direction away from the source.

The model would be considered successful if it climbed the nutrient gradient
and then tended to stay on or near the nutrient source.

II. The Purpose of the Demonstration

The purpose of the demonstration was to prove that a model based on
reinforcement principles _could be constructed_ which would behave as
specified (a proof of principle). It was asserted that one could not.

III. What was Not the Purpose of the Demonstration

It was NOT the purpose of the demonstration to show that

    1. reinforcement-based and PCT models of e. coli behavior are equally
         "valid," robust, explanatory.

    2. reinforcement is a real phenomenon.

To return to our favorite illustration, asserting that NO
reinforcement-based model could behave properly in the test situation is
equivalent to asserting that the Ptolemaic system could not properly
describe the motions of Mars. Ptolemaic theory may be wrong, but is it true
that it can't handle these data?

I am arguing that your assessment of the inability of Ptolemaic theory to
handle the Mars data is incorrect. I am NOT arguing that Ptolemaic theory
is correct! NOR am I arguing that it can handle ALL data.

Big difference, but one that seems to have been missed. You say:

My conclusion is that your e. coli model did not demonstrate that
reinforcement is a real phenomenon; it simply assumed that it was.

Yes, just as my calculations based on cycles and epicycles would not have
demonstrated that cycles, epicycles, and retrograde planetary motion are
real phenomena; they would only have demonstrated what I set out to
demonstrate, that such calculations do provide an account of planetary
motion that is consistent with their observed paths through the night sky.

As I did not set out to demonstrate that reinforcement is a "real"
phenomenon, I cannot be faulted for failing to demonstrate it.

IV. Consistency of ECOLI4a with reinforcement theory.

In attempting to develop a model that would perform according to criterion
in the specified, very unfavorable (with respect to the theory) test
environment, I considered several approaches before arriving at one that
appeared to work. This model incorporated several standard
reinforcement-theory concepts, as described below:

    1. Reinforcement
        This is an increment in the probability of a response when the
        response is immediately followed by a favorable outcome, in this
        case, an increase in the rate of nutrient increase.

    2. Punishment
        This is an decrement in the probability of a response when the
        response is immediately followed by an unfavorable outcome, in
        this case, a decrease in the rate of nutrient increase.

    3. Stimulus Control
        The response probabilities affected by reinforcement/punishment
        are conditional probabilities; the conditions are identifed with
        discriminably different stimulus conditions. Here, we defined a
        rising nutrient gradient as S+ and a falling nutrient gradient as
        S-.

For these relationships to hold, the organim must be assumed to have
appropriate structures that provide the necessary functions. For
reinforcement and punishment to work, there must a sensory structure that
detects the rate of change in nutrient concentration, a structure to store
the rate immediately prior to a tumble, and a structure to compare this rate
to the rate immediately after a tumble. It is not difficult to imagine a
set of molecular components that might provide these functions, but these
would involve pure speculation on my part, so I included in the model only
what they do, not how they do them.
For discrimination to take place, there would also need to be a mechanism
that could selectively associate the stored state of nutrient change prior
to a tumble with the appropriate structural representation of tumble
probability (perhaps the concentration of a chemical whose effect on tumble
probability is mediated by an enzyme whose concentration represents the
stored value of, say S+, but again, such mechanisms are speculative; only
the functions are modeled).

The model works as follows:

    1. When S+ was present prior to a tumble and the tumble is
        reinforced, the probability of a tumble given S+ is increased.

    2. When S- was present prior to a tumble and the tumble is
        reinforced, the probability of a tumble given S- is increased.

    3. When S+ was present prior to a tumble and the tumble is
        punished, the probability of a tumble given S+ is decreased.

    4. When S- was present prior to a tumble and the tumble is
        punished, the probability of a tumble given S- is decreased.

All this is consistent with common principles of reinforcement theory.

V. How the Model Performed in the Test Environment

Because the test environment contained a point source of nutrient, tumbles
made while moving "upstream" (up the nutrient gradient) yielded a more
favorable rate of nutrient change (reinforcement) on only 1/4 of tumbles and
a less favorable rate of nutrient change (punishment) on 3/4 of tumbles.
Over the long run this drove the probability of a tumble in the presence of
S+ down to its minimum value. E. coli virtually stopped tumbling when
moving up the nutrient gradient.

The reverse condition held when e. coli was moving down the nutrient
gradient, so that the probability of a tumble in the presence of S- was
driven upward to its maximum value. E. coli tumbled frequently when moving
down the nutrient gradient.

The result of these two changes in tumble probability was that e. coli
behaves as required by the Challenge. Given the right set of cycles and
epicycles, the Ptolemaic system can indeed reproduce the apparent motions of
Mars.

VI. Replies to Specific Criticisms [Bill Powers (950610.0300 MDT)]

Re: ECOLI3

This model was withdrawn after I realized that I had made some errors of
application. End of discussion.

Re: ECOLI4

When both poles of the switch are vertical, the overall function is
exactly the same as in Ecoli3. In other words, it is predetermined that
delays will be short when going the wrong way, and long when going the
right way. Reaching the maximum effect is slowed, but reaching the
correct effect is inevitable and built in beforehand.

I am having trouble following your diagram but I believe it is wrong. (I
can't make out what those switches are doing, where the pivot is. In
addition some of the critical elements appear to have been mislabeled.) The
mechanism I have given e. coli has no "predetermined," "correct," or
"inevitable" direction of change. It is only the nature of the environment
that causes the two tumble probabilities to move in the "correct" (i.e.,
adaptive) directions.

We went over this in some detail a few months ago and the end result was
that you (finally) got it right. This diagram suggests that you are
confused again.

You may be speaking of the effect of reinforcement and punishment on tumble
probabilities (which is the same whether S+ or S- is present). In that case
your complaint is that I have built the model so that reinforcement
reinforces and punishment punishes. This hardly seems a valid complaint.

The analogous complaint for the control model is that you have given it
negative gain. Why should gain be negative? Why, because if it isn't, e.
coli won't move up the nutrient gradient! Put another way, you assume in
the control model that e. coli "wants" to see nutrient levels increasing. I
assume the same thing in my code for the effects of reinforcement.

The only reason that adding the reinforcement path does not destroy
control completely is that the environmental geometry is centered on a
point-source of nutrient. If the gradient did not converge toward a
point (if it behaved as for a line source or a very distant point
source), the probability of the second derivative being positive would
be 50% regardless of the value of the first derivative. Then the switch
would spend as much time in the wrong position as the right position,
and PTS+ and PTS- would wander at random.

True. So what? The Challenge was not to build a model that would work in
any gradient, it was to build one that would work in the gradient supplied.
This model does. Whether it works in other environments is irrelevant.

Your judgment that the reinforcement model "works" was based only on the
fact that the resulting behavior was correct: E. coli did approach the
target.

That was the criterion for success of the model. As I recall, it was no
trivial exercise to find a model that did "work." But the model had other
criteria as well: it had to be consistent with the principles of
reinforcement theory, and it was.

But I have spent a lot of time in careful analysis of the logic
of your model, trying to understand why it does work, not just that it
does work. And I have found that what makes your model work has nothing
to do with reinforcement theory: adding explicit reinforcement in the
way you did _worsens_ the performance of the model.

I believe I have shown (above) that the model is entirely consistent with
reinforcement theory. It is true that on 1/4 of tumbles the appropriate
conditional probabilities get moved in the "wrong" direction by the model,
given the point source of nutrient. That's the breaks: the model knows
nothing about the effect of its behavior except what it can learn about
these from the appropriate comparisons. Sometimes what it learns on a given
trial is just plain wrong when viewed against the long-term interests of the
bacterium. Because many "trials" give false information, it is important
that the probabilities not be adjusted too much after a given trial: this
would lead to instability. By making the adjustment process move slower
than the process it is adjusting, we prevent instability. I believe you've
pointed to the same necessity when speaking about reorganization, and it's
certainly built into Hans Bloom's adaptive controller. There are some real
lessons you might have learned here about the requirements of an adaptive
system, but you've been too bent on demolishing the model to notice.

When you were arguing that your model did work, you did not go through
the model as I have done to see whether it worked as you said it worked.
You went rapidly through some verbal arguments, but the clincher for you
was that the right result occurred: E. coli approached the target.

Now these are real fighting words, Bill. I suggest you get out those old
posts of mine concerning ECOLI4a and READ THEM CAREFULLY. Talk about
selective memory! Wow!

As I recall, I expended considerable effort carefully describing the
mechanism of ECOLI4a. We went through at least two misdescriptions on your
part and I posted not only a clear diagram of the model's logic but an
equally clear diagram as to how the specific nutrient gradient determined
the outcome of the simulation. Remember those "Marken probabilities," you
know, where Rick said the outcome HAD to be 50-50, so that no learning was
possible, his computer program said so, never mind the diagram? I strongly
encourage you to go review those exchanges and see whether your recollection
of the events matches what appears there.

I think it would be instructive to speculate about why your verbal
arguments seemed sufficient, when in fact they glossed over fundamental
defects in the logic. I think it would be reasonable to say that you
simply couldn't believe that reinforcement theory would not work.
Assuming that it had to work, you didn't see any reason to go through
the details of your system and figure out what it would actually do
according to its own structure, instead of according to what you
expected and wanted it to do. Your reasoning, in fact, was driven by the
goal, being adjusted to make the perception match the goal-perception.

Sorry, Bill. (1) My arguments are sufficient. (2) They do not gloss over
any defects, fundamental or otherwise, in logic. (3) Therefore your
speculations as to why I thought they were sufficient are moot. I thought
they were sufficient because they are sufficient, not because I was being
led by the nose by any forgone conclusions.

Obviously, I don't consider this to be a sin. It is a very common
phenomenon that goes a long way toward explaining the phenomenon of
belief. Belief is not just a passive perceptual phenomenon; it's an
active control process in which inputs are selected that will support
the belief. When we have reason to want a conclusion to be true, we can
construct logical arguments that are just complete enough to support the
desired result, but not so complete as to risk disproving it.

Bill, before you go calling the kettle black, I suggest you take a good,
long look at your own performance in this little debate. You may find it to
be an eye-opener. Even PCT theorists can't escape from behaving as PCT
predicts. (;->

Regards,

Bruce