Conference; Evolutionary "nothing but"

[From Bruce Abbott (970811.2235 EST)]

Bill Powers (970811.0538 MDT) --

Thanks for your excellent summary of the conference proceedings, Bill. I'm
really sorry I couldn't be there to experience it first hand; the research
projects you described sound fabulous. Unfortunately, the meeting
conflicted with my son's public music recital (the final requirement for his
degree in music education at Indiana University).

Bill Powers (970809.2151 MDT) --

Bruce Abbott (970806.1235 EST)

I'm not talking about PCT as Creation Science, by the
way, but rather, the kind of "logic" implicit in the series of statements by
you and Rick that were quoted in my post.

Rick and I don't coordinate our answers to email; it's a bit disconcerting
to find my posts interleaved with Rick's to generate a nonsensical series
of citations.

I agree that the problem is not with your posts. However, in my response,
your posts were not interleaved with Rick's. Your statements from two posts
(given in order) were followed by Rick's response in his post. If the
result is nonsense, I had nothing to do with that. I found it just as
disconcerting as you do.

Regards,

Bruce

[From Bill Powers (970812.0555 MDT)]

Bruce Abbott (970811.2235 EST)--

Bruce Abbott said (970806.1235 EST)

I'm not talking about PCT as Creation Science, by the
way, but rather, the kind of "logic" implicit in the series of statements by
you and Rick that were quoted in my post.

And I replied

Rick and I don't coordinate our answers to email; it's a bit disconcerting
to find my posts interleaved with Rick's to generate a nonsensical series
of citations.

To which you responded
I agree that the problem is not with your posts. However, in my response,
your posts were not interleaved with Rick's. Your statements from two posts
(given in order) were followed by Rick's response in his post. If the
result is nonsense, I had nothing to do with that. I found it just as
disconcerting as you do.

···

------------------------------------
The interleaving was done in your mind, allowing you to combine my
statements and Rick's as if they were part of a single argument:
---------------------------------
You said:
Now this is a very interesting line of argumentation. I am told that I
cannot establish the reality of negative feedback in Darwinian evolution
through _verbal_ argument[Bill's statement], and that any _mathematical
derivation_ I might come up with would be meaningless, because I could
"probably make it come out the way [I] like." [Rick's statement]
Congratulations, gentlemen: you've created a _perfect_ Catch 22. I'm
damned if I do, and damned if I don't.

[Note that you are forming a logical AND in which one proposition is mine
and the other is Rick's]

Unfortunately for you, the argument is a two-edged sword. You claim that e.
coli is a negative feedback system? You can't establish that through verbal
argument (according to your own rules), and even if you could demonstrate it
mathematically (which so for you have been unable to do), why, you've just
made the analysis come out the way you like (according to Marken).

I'm sorry, gentlemen, but this is Creation Science.
--------------------------------------------------------
The "you" throughout is clearly plural, as if Rick and I are making one
argument against yours: you address your remarks to "gentlemen."

I made it quite clear in my earlier posts that I was attempting a
mathematical analysis of the E. coli situation without any guarantee that I
could do it just then, and when I failed I said so. I said that I don't
think that natural selection involves negative feedback, and that negative
feedback is often read into situations where it doesn't really exist; I
thought I had given an illustration of that with the mass on a spring. I
also said that I can't prove that my feeling about natural selection is
right. Later, I did find a valid analysis showing negative feedback in the
E. coli case, using a somewhat different model, and I also explained why my
initial model gave the reference signal no effect. You haven't replied to
any of that. Neither have you shown any reason for anyone to believe -- yet
-- that natural selection involves negative feedback. You have simply
insisted that it does.

Rick has a way of throwing his conclusions onto the table without bothering
to explain how he justifies them. If that gets him into trouble with you,
that's his business and not my problem. Don't try to make it mine. If you
and Rick could get off the plane of "tis so, taint so" your squabbles would
be more interesting. As it is, I would just as soon be left out of them.
--------------------------------------------
That said, I do agree with Rick that operant-behavior analysts use a style
of mathematical analysis that can be made to reach almost any conclusion.
There is never any single coherent progression of mathematical derivations;
instead, the derivations are frequently carried over sticky places with
verbal arm-waving, and mathematical identities are frequently confused with
solutions of equations. The conceptual biases of the analysts are built
into the initial assumptions they make -- just look at how often the
authors of a paper will refer to "stimulus control" when nothing they have
said supports that concept in any way. Look at Herrnstein's mathematical
treatment of the matching law, which any engineer would recognize as a
simple statement that the two schedules are identical. In fact, the
equations are valid mathematical statements ONLY if the two schedules are
actually identical.

The whole idea of a mathematical analysis is to leave yourself as little
room as possible to influence the outcome because of beliefs, informal
interpretations, and unspoken assumptions. Implicit in a mathematical
analysis is an agreement to abide by the results, even if they don't
coincide with what you thought or hoped they would be. This doesn't happen
in the literature I have seen. Instead, an analysis that by itself would
lead to the "wrong" interpretation is bolstered by as many additional
assumptions as necessary to support the party line, and the only reason for
accepting the added assumptions is precisely to make the conclusions come
out right. If you want to draw a parallel with the reasoning used in
creation science, this kind of analysis would provide an excellent
opportunity. It's my opinion that self-styled behavior analysts don't
understand systems analysis at all. They just mess around with equations
until they can interpret the results in a familiar way. I haven't yet seen
anything I would recognize as an objective analysis in that genre.
-----------------------------------
While we're on this subject, I think you need to reconsider your own E.
coli model, which you seem to remember as having settled the issue of
reinforcement theory as far as its justification is concerned. You may
recall (or you may not, as you have never commented on this), your model
involved four cases, two of which increased the probability-density of a
response to an increase or decrease in the critical variable, and two of
which decreased it. Together, these four cases resulted in a slight
preponderance of changes in delay time in the right direction. However, two
of the cases worked against the other two: they went in the wrong
direction. So your model actually contained a conflict.

I suggested that it may have been the geometry of the particular situation,
in which the gradient converged toward a point, that created enough
imbalance to give the right result. I suggested that there might be other
geometries in which the balance would go the wrong way. If I had been the
one offering your model, and had heard such suggestions, I wouldn't have
rested until I had shown _one way or the other_ whether these allegations
were true. But you seemed totally uninterested in pursuing that matter
further.

If you will remember, your model began with the observation that if the
system were moving the wrong way, the probability that a tumble would
improve the situation was greater than 0.5, and if it were moving the right
way, the probability of a tumble making the situation worse was also
greater than 0.5. So clearly what was needed was a way of reinforcing
tumbles that were responses to a decrease in the critical variable, and
reinforcing _lack_ of tumbles that were responses to an increase in that
variable. Logic also required that a tumble in response to an increase in
the critical variable had to be punished, and _lack_ of a tumble in
response to going the wrong way had to be punished.

You built this logic into your model. To support this logic, you had to
assume that many detailed operations were being carried out inside the
organism: there had to be comparison of present circumstances with past
results of tumbles, and of present values of the critical variable with
past values. In short, you built into the model exactly the observations
you had made about the relationships of probabilities of improving the
situation to the effects of tumbles in the various cases. And you built
into it your own logic by which you had arrived at this understanding.

At best, the result was not a model of organisms in general, but only a
model of organisms capable of grasping the situation and applying logic to
it in just the way you did. You modeled an organism capable of logical
thinking.

As it worked out, the correct behavior was produced not by rewarding the
right changes in behavior, but by punishing the wrong changes in behavior.
There was a provision for rewarding the right change in behavior, _but it
worked the wrong way_. Fortunately, punishing the wrong change in behavior
had a somewhat larger effect on the outcome than rewarding the right change
in behavior, so -- by luck -- the overall model exhibited the right behavior.

I laid out all these problems in our original discussions, but all you have
chosen to remember is that you got the right result. If I had been in your
shoes, I would have been very concerned about these internal
contradictions, and I would have wondered if there might be situations in
which this particular arrangement of countervailing effects might come out
wrong. But I didn't see any interest on your part in pursuing that matter
any further: you got lucky and weren't about to question the result. Your
intention was to show that reinforcement theory was capable of handling the
situation just as well as a control model could, and apparently, as far as
you were concerned, you had achieved your objective. Why rock the boat?

If you step back a bit and look at what you demonstrated with your model,
you will see that you argued directly against the idea that stimuli have
reinforcing properties. You modeled an organism capable of logical
reasoning, and by that means you showed that the correct result could be
obtained, in the particular situation imagined, through the medium of
logic, and not through any "reinforcing" effect of anything in the
environment. By attempting to defend the concept of reinforcement, you
succeeded, if only in one special case, in destroying it.

Best,

Bill P.

[From Bruce Abbott (970812.1845 EST)]

Bill Powers (970812.0555 MDT) --

The interleaving was done in your mind, allowing you to combine my
statements and Rick's as if they were part of a single argument:

Interleaving would involve reordering your statements to alternate with
Rick's, not presenting them in their actual order (Rick's after yours) as I
did. "Concatenating" is the term you want.

[Note that you are forming a logical AND in which one proposition is mine
and the other is Rick's]

Yes, that is what I did.

Rick has a way of throwing his conclusions onto the table without bothering
to explain how he justifies them. If that gets him into trouble with you,
that's his business and not my problem.

But it _is_ your problem if you do not voice your disagreement with those
conclusions. Most readers of CSGnet assume that you and Rick are of like
mind on most issues relating to theory, unless you offer a correction (as
you sometimes do). There being no such correction offered here, Rick's
addition will be accepted as if you were in complete accord with it. And
if these are the rules by which my view is to be judged, then in effect you
and Rick between you have ruled out the possibility of proof, mathematical
or otherwise. You don't see anything wrong with this?

That said, I do agree with Rick that operant-behavior analysts use a style
of mathematical analysis that can be made to reach almost any conclusion.

But we weren't discussing using "a style of mathematical analysis that can
be made to reach almost any conclusion." We were talking about analyzing a
specific model using the same style of mathematical analysis used by you and
others to establish whether a particular closed-loop system is or is not a
negative feedback system. If Rick meant some other style of mathematical
analysis, he should have said so. In the context in which it appeared, I
could only assume that he was talking about the sort of mathematical
analysis you were undertaking for e. coli. In that context, what would be
the point of switching the subject?

Given the type of analysis we were actually discussing, the rest of your
complaint about mathematical analysis as carried out by some EAB types is
irrelevant to the present discussion, although I would like to point out for
the record that you have _once again_ returned to your old myth about
Herrnstein's matching law. How many times do I have to go over this with
you? Herrnstein's matching law was formulated to describe performance on
concurrent VI VI schedules (in which case it is most definitely NOT a simple
statement that the two schedules are identical), whereas your assertion
would be correct _only_ if it were formulated to apply to concurrent ratio
schedules, which it wasn't. Apparently the charge that "any engineer" would
see the flaw is just too delicious for you to resist. Nevermind the fact
that it isn't true.

While we're on this subject, I think you need to reconsider your own E.
coli model, which you seem to remember as having settled the issue of
reinforcement theory as far as its justification is concerned. . . .

No, no, no, no, no. I never claimed any such thing. You and Rick claimed
that no reinforcement analysis could account for your simulated e. coli's
behavior; I demonstrated that one could. That is all.

You may
recall (or you may not, as you have never commented on this), your model
involved four cases, two of which increased the probability-density of a
response to an increase or decrease in the critical variable, and two of
which decreased it. Together, these four cases resulted in a slight
preponderance of changes in delay time in the right direction. However, two
of the cases worked against the other two: they went in the wrong
direction. So your model actually contained a conflict.

After two years I'm getting fuzzy about the details, but I believe it was I
who pointed this out to you, not vice versa, in one or more of my repeated
attempts to get across its logic to you and Rick. Both of you kept telling
me that the thing wouldn't work as advertised, and Rick even went so far as
to write a computer program that "proved" that the model would do no better
than a random walk. It was a nightmare, but I eventually succeeded in
getting across to you _that_ the system worked, and that it worked _as I had
described_.

I suggested that it may have been the geometry of the particular situation,
in which the gradient converged toward a point, that created enough
imbalance to give the right result. I suggested that there might be other
geometries in which the balance would go the wrong way. If I had been the
one offering your model, and had heard such suggestions, I wouldn't have
rested until I had shown _one way or the other_ whether these allegations
were true. But you seemed totally uninterested in pursuing that matter
further.

No, I believe it was I who pointed out that the model took advantage of the
geometry of the situation, and that it probably wouldn't work in some other
geometries. This business about me being uninterested in "your" allegations
and uninterested in pursuing the matter is a figment of your imagination.
There was no need for me to do so, as I had already freely admitted that
this was true.

If you will remember, your model began with the observation that if the
system were moving the wrong way, the probability that a tumble would
improve the situation was greater than 0.5, and if it were moving the right
way, the probability of a tumble making the situation worse was also
greater than 0.5. So clearly what was needed was a way of reinforcing
tumbles that were responses to a decrease in the critical variable, and
reinforcing _lack_ of tumbles that were responses to an increase in that
variable. Logic also required that a tumble in response to an increase in
the critical variable had to be punished, and _lack_ of a tumble in
response to going the wrong way had to be punished.

What was needed was a way to reinforce tumbles when the organism was going
the wrong way (away from the nutrient source) and to punish tumbles when the
organism was going in the right way (toward the nutrient source); there was
no provision for reinforcing or punishing lack of tumbles. Lack of a tumble
is not an event.

You built this logic into your model. To support this logic, you had to
assume that many detailed operations were being carried out inside the
organism: there had to be comparison of present circumstances with past
results of tumbles, and of present values of the critical variable with
past values. In short, you built into the model exactly the observations
you had made about the relationships of probabilities of improving the
situation to the effects of tumbles in the various cases. And you built
into it your own logic by which you had arrived at this understanding.

Let us imagine that I had challenged you to construct a control-system model
that would produce a certain behavior. I claim that it is impossible for
you to do so for this particular situation. It would be up to you to
specify the controlled variable and the detailed structure of the model
(e.g., single-level, two-level, proportional, proportional plus derivative,
proportional plus integral, and so on). To my surprise, you come up with a
model that exhibits the required behavior. Now I come back and cry "foul"
-- you specified just what was needed to make the model behave the way you
wanted it to! Would you accept this criticism as valid? I don't think so.

Remember, we were not talking about real e. coli. The task was not to
develop a realistic reinforcement model of e. coli (absurd, as there is no
evidence that e. coli learns to control nutrient level, and the
reinforcement model is a learning model) but to develop a model consistent
with reinforcement principles that would perform as required. Under these
conditions, I was free to invent, so long as the resulting system was so
consistent. If I had been trying to develop a model to account for real
experimental data, my modeling would have been constrained by observation;
I would have needed to demonstrate that the variables supposedly being
sensed and acted on were actually sensed, that the putative reinforcer
really was acting as the reinforcer, and so on. But this was not the case.
I had no such constraints, and therefore was free to suppose whatever I
needed to suppose to make the model work, so long as the result did not
violate basic reinforcement principles.

At best, the result was not a model of organisms in general, but only a
model of organisms capable of grasping the situation and applying logic to
it in just the way you did. You modeled an organism capable of logical
thinking.

Not really. I could easily make up a story in which simple biochemical
variables interacted so as to realize my hypothetical mechanism. But this
would be silly, as from the beginning I stated that the model was not
intended to be realistic. Remember, its purpose was only to show that a
model consistent with reinforcement theory would behave properly under the
conditions given.

As it worked out, the correct behavior was produced not by rewarding the
right changes in behavior, but by punishing the wrong changes in behavior.
There was a provision for rewarding the right change in behavior, _but it
worked the wrong way_. Fortunately, punishing the wrong change in behavior
had a somewhat larger effect on the outcome than rewarding the right change
in behavior, so -- by luck -- the overall model exhibited the right behavior.

There was no luck involved (the model converges on the right behavior,
homing in on the source of nutrients, rapidly and in every run of the
program). Furthermore, your description is grossly inaccurate. The "right"
behaviors (tumbling when going in the "wrong" direction, away from the
nutrient) were sometimes rewarded and sometimes punished; similarly, the
"wrong" behaviors (tumbling when going in the "right" direction, toward the
nutrient) were sometimes rewarded and sometimes punished. Because of the
circular geometry of the nutrient density field, right behaviors were
reinforced more often than they were punished, and wrong behaviors were
punished more often than they were reinforced. As a result, the
probabilities of the right behaviors quickly increased to a high value, and
the probabilities of the wrong behaviors quickly diminished to a low value.
Your suggestion that punishment for wrong "change in" behavior somehow
"outweighed" rewarding the right "change in" behavior is not only incorrect,
it isn't even logical. If your description were true, the model should have
perversely learned to do exactly the wrong thing.

I laid out all these problems in our original discussions, but all you have
chosen to remember is that you got the right result. If I had been in your
shoes, I would have been very concerned about these internal
contradictions, and I would have wondered if there might be situations in
which this particular arrangement of countervailing effects might come out
wrong. But I didn't see any interest on your part in pursuing that matter
any further: you got lucky and weren't about to question the result. Your
intention was to show that reinforcement theory was capable of handling the
situation just as well as a control model could, and apparently, as far as
you were concerned, you had achieved your objective. Why rock the boat?

Again, this is a total fantasy on your part. It didn't happen that way.

What _has_ happened is that you wish to demonstrate that my thinking is of
just the sort you attributed to those "operant behavior analysts" whose
faulty analytical methods were described earlier in your post. So you've
found a way to "remember" those past discussions in just the right way, so
that the parallel will be obvious to all.

Your description of our interchanges on the issue is in part a wishful
reconstruction of the facts; I stand on the record (CSGnet archives). If
you don't believe me, you can look it up.

This is the second time I have challenged you to do this; the first time,
you took up that challenge and discovered, much to your chagrin, that you
were wrong, that my description of our interchanges was accurate. You said
it was a lesson you would not soon forget. Have you forgotten it already?

Returning to our _present_ discussion, what argument would satisfy you that
Darwinian evolution involves a negative feedback loop? Bear in mind that
the argument I present will have to be verbal; I am not a mathematician.

Regards,

Bruce

[From Bill Powers (970812.2212 MDT)]

Bruce Abbott (970812.1845 EST)--

It's pointless to argue about who said what when if we don't dig up the
archives and document the claims. Better to go to the source and see what
we're talking about.

Here is the description of one of your operant models:

* Created: 11/14/94 *
* *
* This program implements a discriminated operant simulation of the *
* "tumble-and-swim" behavior of an "e. coli" capable of learning from *
* experience. If a tumble in the presence of S+ (rising nutrient *
* level) results in a more positive rate of nutrient change, *
* p(Tumble|S+) increases, otherwise it decreases. If a tumble in the *
* presence of S- (decreasing or steady nutrient level) results in a *
* more positive rate of nutrient increase, then p(Tumble|S- increases, *
* otherwise it decreases. Thus the effect of the change in the rate *
* of nutrient change following a tumble on p(Tumble|S+) and *
* p(Tumble|S-) is symmetrical. Experience with the consequences *
* of tumbling gradually "shapes" e. coli's behavior so as to maximize *
* nutrient levels. *

Look at the actual phenomenon being modeled: what is required is that a
tumble occur sooner if the nutrient level is decreasing, and later if it is
increasing. That is what the model has to accomplish, because that is the
only way that progress up the gradient can be achieved.

An increase in delay corresponds to a reduction in the probability of a
tumble per unit time (or per iteration); a decrease in delay is equivalent
to an increase in that probability. So the output function of the "operant"
model is just the inverse of the output function of the PCT model.

The reinforcer is an increase in the rate of change of nutrient across a
tumble: Dnut - Nutsave, which you call DeltaNutRate, or DNR for short. The
discriminative stimulus is Nutsave itself, the rate of change of nutrient
prior to the previous tumble. The output probabilities are called
pTumbleGivenS+ and pTumbleGivenS-.

Your program sets up a four-way decision:

If DNR > 0 then {reinforcement}
   If nutsave > 0 then [increase pTumbleGivenS+]
   else [increase pTumbleGivenS-]
else {punishment}
   if nutsave <= 0 then [decrease pTumbleGivenS+]
   else [decrease pTumbleGivenS-]

Remember that what is required is for pTumbleGivenS- to increase and
pTumbleGivenS+ to decrease. This means that the two cases which work in the
right direction are the statements containing [increase pTumbleGivenS-] and
[decrease pTumbleGivenS+]. Those two statements, in fact, are the PCT model
of E. coli's method of swimming up a gradient: if the direction is up the
gradient (S+), lengthen the time to the next tumble (decrease the
probability of a tumble per iteration), and the opposite for directions
down the gradient.

The OTHER two possibilties, which occur whenever DNR is positive and
preceded by a positive value of Dnut, or negative and preceded by a
negative value of dNut, produce an increase in pTumbleGivenS+ or a decrease
in pTumbleGivenS-, both of which are in the wrong direction. There should
never be an increase in pTumbleGivenS+, or a decrease in pTumbleGivenS-. In
order for the right overall result to occur, even at reduced efficiency,
it would be necessary for decreases in pTumbleGivenS+ to exceed increases
in pTumbleGivenS+, and for increases in pTumblegivenS- to exceed decreases
in pTumbleGivenS-.

This means that the change in direction across a tumble must such that when
dNut is positive, the next value of dNut is more than 50% likely to be less
positive. This is indeed the case, as is the opposite case for dNut negative.

My initial judgement was based on the idea that the wrong cases were just
as likely as the right cases. I then worked out the actual probabilities
given the geometry, and realized that I had been mistaken: the geometry was
such that the required preponderance of the "right" cases existed. I duly
announced this correction of my mistake.

The main operational difference between the operant and the PCT models is
simply that the PCT model never changes the delay in the wrong direction.

In fact, the only required condition is that pTumbleGivenS+ be small, and
pTumbleGivenS- be large. Introducing the previous state of dNut and the
value of DeltaNutRate only reduces the speed of progress up the gradient in
comparison with the PCT model, by assuring that on some tumbles, the
probabilities (and hence the delay before the next tumble) will change the
wrong way. I believe I calculated that 3/8 of the probability changes would
be in the wrong direction, although I can't vouch for that number. If that
number were right, it would mean that the operant model would proceed up
the gradient 1/4 as fast as the PCT model.

Your model exhibits "learning" only because the probabilities -- in both
the right and wrong directions -- increase with time. The PCT model could
be given the same "learning" capability just by increasing with time the
constant that converts error into delay.

Note that NEITHER model could learn to swim _down_ the gradient of a
_repellent_ without a change in the program. There is only one outcome
possible, and it is due strictly to the changes in net probabilities toward
higher values of one and lower values of the other. If these probabilities
were set, respectively, to 1 and 0 initially, the same result would occur
immediately that takes time to develop when the probabilities (or delays)
are changed gradually.

···

----------------------------------------------------------------------------
That is my reconstruction of my previous analyses of your model, including
my mistake.
-----------------------------------------------------------------------------
Your model is based on very general propositions, as given in your comments
above: a relationship between reinforcement given S+ or S- and changes in
probabilities. In fact, your definitions embody the essence of
reinforcement theory, as I understand them.

So consider this:"If a tumble in the presence of S+ (rising nutrient level)
results in a more positive rate of nutrient change, p(Tumble|S+) increases,
otherwise it decreases."

If that is really the general statement, we have just the wrong
relationship for E. coli in the first part of the statement. If dNut
increases as a result of a tumble, and dnut is positive, this statement
says that the probability of a tumble should _increase_. In fact, it should
decrease. So in your description of the basic rule, you are actually
describing one of the "wrong" cases first!

The case that works in the right direction is this: if there is an increase
in the rate of change of nutrient across a tumble, then the probability of
a tumble in the presence of a negative rate of change of nutrient should
increase. This is the case where the change in dnut is positive but the
value of dnut is still negative; this should result in an increase in the
probability of a tumble in the presence of S-.

Similarly, your second statement says, "If a tumble in the presence of S-
(decreasing or steady nutrient level) results in a more positive rate of
nutrient increase, then p(Tumble|S- increases, otherwise it decreases."
Since we _always_ want pTumble|S- to increase, the "otherwise" part of this
statement works in the wrong direction.

In fact, the parts of these statements that refer to reinforcement are
logically irrelevant. The only two statements that are needed are

1. If the nutrient rate of change is positive (S+), decrease the
probability of a tumble given S+, and

2. If the nutrient rate of change is negative (S-), increase the
probability of a tumble given S-.

This will achieve the same result and (if I remember my old analysis
correctly) do it four times as fast as when "reinforcement" is included as
a condition. Including the reinforcement condition _slows the learning_
without adding any extra generality.

The above two statements are just the PCT model expressed in terms of
changing probabilities. This model incorporates "learning" simply because
the output operation includes an appropriate change in probability, as in
the operant model.

Pause for comment.

Best,

Bill P.