Making Progress

[From Bruce Abbott (941130.1915 EST)]

Bill Powers (941130.0725 MST)

Bill-- I've tried twice now to respond to your earlier post concerning a
traditional reinforcement theory analysis of the tracking experiments. The
basic thrust is that I think it might be an interesting exercise, but I'm not
sure this is something I want to do at this time. Although my background is
strongly Skinnerian, my views are not. (I've cited Skinner more often in CSG-L
than previously in my entire life.)

I'm more interested in attempting to apply and evaluate PCT in a field
traditionally dominated by reinforcement theory than in playing the role of
Resident Behaviorist. Partly this means rethinking many issues and concepts
from a PCT perspective that are dealt with in very different ways from a TRT
viewpoint. My "defense" of the law of effect is one result: I recognized a
principle at work that did not seem (to me at least) to contradict PCT,
although it clearly needed to be reformulated from that perspective. In so
doing I tried to differentiate between the empirical law of effect, which is
just a description of what one can see happening, and the theoretical law of
effect, which includes unobserved entities like associative connections. I
have been arguing for the former and not the latter.

Therefore I am heartened that we seem to be coming to some kind of agreement
[I'm not sure that we ever did disagree, but it sure felt like it from here],
as indicated in the following:

There is
nothing basically wrong with what the Law of Effect proposes: Stated
from the PCT perspective, the proposal is that we (all organisms that
learn) learn to control by finding actions that have systematic effects
on perceptions that are important to us. Even Albert Einstein had to
learn which way to twist the cap on his tube of toothpaste to get it
open.

More than that, we learn to perform those actions that bring these perceptions
under control. But then there is the question of motivation:

PCT, however, introduces a new consideration: the question of whether
Albert wanted to brush his teeth with toothpaste. If he wanted to use
baking soda, or didn't use anything but water, or didn't want to brush
his teeth at all, the direction of effort required to take the cap off
the toothpaste, while still remaining a fact, would be irrelevant. What
gives the Law of Effect its lawful-seeming force is not any particular
consequence, but the fact the organism intends for certain consequences
to occur.

Thorndike was acutely aware of this problem and tried to resolve it. Why
should behavior change simply because certain activities are followed
systematically by certain "states of affairs"? He concluded that the answer
lay partly inside the organism, but that given the current state of ignorance
(in 1911) one could only speculate as to why some consequences affected the
likelihoods of the behaviors that produced them. He noted that biological
need was not a sufficient answer because there were many things for which
humans and other animals would work that were actually bad for them, such as
consuming too many sweets or taking addictive drugs. So he settled for an
empirical test: observe whether the person or animal willingly approaches a
given state of affairs, does or does not resist being brought into that state
of affairs, or actively attempts to withdraw from it.

In observing these effects, Thorndike came tantalizingly close, in my view, to
discovering the mechanism of feedback regulated control. A satisfying state
of affairs is some state OF THE ORGANISM that the organism will act to attain,
if necessary, and will then resist being removed from: the relationship
between the state and behavior is one of negative feedback. An annoying state
of affairs is some state of the organism that the organism will act to avoid,
if necessary, and will actively remove itself from; again the relationship is
one of negative feedback, but now the state to be attained is the ABSENCE of
the state-of-affairs being tested. Only a small--but conceptually difficult--
step separates these observations from their explanation in terms of feedback-
regulated control.

Thorndike's law of effect, then, asserts that those activities taking place
within a given situation will tend to be repeated if they bring about a
satisfying state of affairs or prevent or eliminate an annoying state of
affairs. Put another way, behavior that tends to reduce error in a perceptual
control system will be acquired and incorporated into the behavioral output
function of the system, whereas behavior that tends to increase error in a
perceptual control system will tend to be inhibited in the behavioral output
function.

So how does the organism discover which behaviors increase or decrease error
in a perceptual control system, and how do these discoveries lead to altered
performance of the behavioral output sub-system? The answer (or more likely,
answers) to this question depends partly on the resources that can be brought
to bear on the problem. E. coli has no resources at all; it must depend on an
innate mechanism to supply the right behavior from the start. Thorndike's
cats appear to lack the capacity to "think it through" and so probably depend
mainly upon a relatively simple mechanism involving something very close to
Thorndike's "selecting and connecting," a mechanism that results in highly
stereotyped behavior that only gradually drifts into more efficient modes of
action. Humans may use this mechanism also, but can also employ other means
such as imitation, imagination, and symbolic manipulation to identify
potential behavioral "solutions."

Thorndike himself suggested that the behaviors appearing in his puzzle boxes
were probably a complex mix of instinctively-driven responses (ready-made
solutions that have worked in similar situations in the organism's
evolutionary history) and previously learned strategies acquired in similar
situations by these street-wise former denizens of the alley. From this mix
would be selected those behaviors that proved effective in reducing those
errors in the cat's perceptual control systems that had been induced by the
disturbances of confinement and lack of access to the tasty treats provided
outside the box.

As to the specific "law of effect" mechanism built into ECOLI4a, it was only
intended to illustrate the principle, not to offer a serious proposal of a
specific mechanism. I agree with you that one should strive whenever possible
to provide mechanisms whose proposed components might be found to be
represented in some real form within the actual organism being modeled, as in
the original e. coli simulation. It is for this reason that the tests between
my specific implementation of a "law of effect" model and the "PCT" model seem
to me to be beside the point. On the other hand, they do very nicely
illustrate how one goes about evaluating the relative merits of competing
models.

I think it important to explore the learning mechanisms involved in these
adaptive changes in behavior, not only because they are interesting in and of
themselves, but because the evolution of such behaviors within the context of
a particular experimental situation is often explained by and expected from a
traditional reinforcement analysis. An analysis that deals only with the
steady state and is silent on how that state was achieved will not
successfully compete with one that appears to do both (even if appearances are
deceiving.)

As a start in this direction, how do you define learning, of the type seen as
an organism constructs an appropriate behavioral output sub-system that brings
a perceptual variable important to it under control? It is only by having an
explicit definition that we can agree whether a particular model does or does
not learn.

Regards,

Bruce