Reinforcement

[philip 4/5/15 11:30am]

I read a book by Karen Pryor (a very famous trainer) about training animals using a method of operant training known as “clicker training”. She emphasized that the most effective way to train an animal was to use a clicker to signal the correct behavior and to avoid using any form of negative reinforcement (punishment) at all (which is often relied upon heavily for obedience training). All undesired behavior is simply ignored by the trainer, and the animal is never explicitly told to avoid a particular behavior. Basically, the clicker is used every time to signal the correct behavior, and food rewards are provided copiously and all at once at the end of the training session. Thus, the food rewards are not directly used in shaping the animal’s response, because the rate of food intake is not a controllable parameter. The clicker training method allowed the process of training to have nothing to do with deprivation and subsequent rewarding. Animals trained using a clicker were healthier, learned tricks more easily, and could even be made to improvise on cue.

I think the take-home message of the book was that the clicker was key for the animal to understand how to behave, and why to behave that way. The other important thing was to focus on establishing an ideal training environment, in which there are

  1. no consequences to fear (no chance of punishment/negative-reinforcement)
  2. very simple, highly predictable parameters (in the form of the clicks) for the animal to seek from the trainer.

I believe there is much to learn from this purified branch of behaviorism, and I suggest that operant training (rather than conditioning) be understood as a very positive thing. There should be room for an analysis of a negative-feedback-only system in a positive-feedback-only environment. Furthermore, I think it’s interesting if PCT incorporated this theory of animal training, because it takes the emphasis off the PHYSICAL disturbance as the cause of the behavior (learning). This puts the emphasis on the nature of the reinforcement rather than on the nature of the disturbance, as far as learning is concerned. And I think learning should be more associated with a process of reinforcement, rather than as a response to a disturbance.

···

[From Rick Marken (2015.04.04.1730)]

···

philip (4/5/15 11:30am) –

PY: I read a book by Karen Pryor (a very famous trainer) about training animals using a method of operant training known as “clicker training”.

RM: This is very interesting. Could you give a little more detail on how the clicker training works?

PY: She emphasized that the most effective way to train an animal was to use a clicker to signal the correct behavior and to avoid using any form of negative reinforcement (punishment) at all (which is often relied upon heavily for obedience training).

RM: For example, how does the clicker “signal” the correct behavior?

PY: All undesired behavior is simply ignored by the trainer, and the animal is never explicitly told to avoid a particular behavior.

RM: Yes, I think that’s also the advice that would be given by any operant conditioner.

PY: Basically, the clicker is used every time to signal the correct behavior, and food rewards are provided copiously and all at once at the end of the training session.

RM: Then I don’t see how the clicker works. Does the animal have to do the same thing after each click in order to get the food at the end of the session? How does it know what to do when the clicker clicks?

PY: Thus, the food rewards are not directly used in shaping the animal’s response, because the rate of food intake is not a controllable parameter.

RM: So the food has nothing to do with the training? I don’t get it.

PY: The clicker training method allowed the process of training to have nothing to do with deprivation and subsequent rewarding. Animals trained using a clicker were healthier, learned tricks more easily, and could even be made to improvise on cue.

RM: I would really like to know how this method works. And it would be great if you could give us an analysis of what’s going on from a control theory perspective.

PY: I think the take-home message of the book was that the clicker was key for the animal to understand how to behave, and why to behave that way. The other important thing was to focus on establishing an ideal training environment, in which there are

  1. no consequences to fear (no chance of punishment/negative-reinforcement)
  2. very simple, highly predictable parameters (in the form of the clicks) for the animal to seek from the trainer.

PY: I believe there is much to learn from this purified branch of behaviorism

RM: I agree. I think it would be really nice to have an anlysis of this training technique from a control theory perspective.

PY: , and I suggest that operant training (rather than conditioning) be understood as a very positive thing. There should be room for an analysis of a negative-feedback-only system in a positive-feedback-only environment. Furthermore, I think it’s interesting if PCT incorporated this theory of animal training,

RM: What do you think should be incorporated into PCT? I’d be surprised if PCT doesn’t already explain how and why her training system works.

PY: because it takes the emphasis off the PHYSICAL disturbance as the cause of the behavior (learning).

RM: I think if you provided a PCT analysis of the training technique it would help me understand what you mean when you say that the technique takes the emphasis off the disturbance as a cause of behavior. Disturbing a controlled variable is not the only way to control behavior. And in the case of learning new behavior, as is the case in animal training, disturbance resistance can’t really work at all since you want the animal to develop a new control organization and disturbance to a controlled variable only works with existing control organizations.

PY: This puts the emphasis on the nature of the reinforcement rather than on the nature of the disturbance, as far as learning is concerned. And I think learning should be more associated with a process of reinforcement, rather than as a response to a disturbance.

RM: Actually, using reinforcement as the basis of training is somewhat more coercive than using disturbance resistance. In order to use disturbance resistance to control behavior the animal already has to have very good control of the variable that is going to be disturbed; indeed, the better the animal can control a variable, the better you can control the actions the animal uses to protect the variable from your disturbances. Using reinforcement requires that the trainer become the only source of the reinforcement so that the reinforcement is provided only when the animal does the desired (by the trainer) behavior. But maybe Karen does something different than what is done in operant learning. If so, I’d be interested in what it is.

Best

Rick


Richard S. Marken, Ph.D.
Author of Doing Research on Purpose.
Now available from Amazon or Barnes & Noble

[From Bruce Abbott (2015.04.08.1030 EST)]

Philip Yeranostan’s description of clicker training is a bit off the mark. Clicker training uses the clicker to reinforce the behavior that the click followed. The clicker is known as a “bridging stimulus,� because it bridges the time between that act and the delivery of the primary reinforcer, allowing the animal to associate that reinforcer with the successful act.

The bridging stimulus is said to function as a conditioned reinforcer because of its association with the primary reinforcer, usually food. As such it provides immediate reinforcement for the act. This is especially important in situations where the primary reinforcer cannot be delivered immediately after the act. For example, in dolphin training (in which a fish is given as the primary reinforcer), it usually is not possible to deliver the fish immediately after the act that the trainer wants to reinforce, as when the dolphin does a back-flip and is now underwater and 10 yards away. The click marks the back-flip as the successful act and the dolphin then approaches the trainer to get its reward.

The click resolves the dolphin’s “assignment of credit� problem, i.e., which of its behaviors should it credit with the earning of the reward? Without the clicker and with the inevitable delays between action and primary reinforcer delivery, it can be very difficult to pinpoint what act resulted in reinforcer delivery. The clicker makes it clear.

Â

Below are two YouTube links to clicker training examples with cats. The first demonstrates the technique in the earliest stages of training. The second demonstrates a variety of tricks that a cat has been taught using the technique. (Note that in the second video the “clicker� is actually a device that plays musical notes.)

https://www.youtube.com/watch?v=t6gRx-a8j4A

https://www.youtube.com/watch?v=puYANVYxPys

By the way, Philip used the terms “punishment� and “negative reinforcement� as equivalent. In fact they are opposites. Punishment suppresses behavior; negative reinforcement reinforces it (through the prevention or removal of an aversive stimulus).  As Philip correctly notes, neither is used in clicker training.

I leave it as an exercise for the reader to explain these results in PCT terms.

Bruce

Rick Marken (2015.04.04.1730) –

philip (4/5/15 11:30am)

PY: I read a book by Karen Pryor (a very famous trainer) about training animals using a method of operant training known as “clicker training”.

RM: This is very interesting. Could you give a little more detail on how the clicker training works?

PY: She emphasized that the most effective way to train an animal was to use a clicker to signal the correct behavior and to avoid using any form of negative reinforcement (punishment) at all (which is often relied upon heavily for obedience training).

RM: For example, how does the clicker “signal” the correct behavior?

PY: All undesired behavior is simply ignored by the trainer, and the animal is never explicitly told to avoid a particular behavior.

RM: Yes, I think that’s also the advice that would be given by any operant conditioner.

PY: Basically, the clicker is used every time to signal the correct behavior, and food rewards are provided copiously and all at once at the end of the training session.

RM: Then I don’t see how the clicker works. Does the animal have to do the same thing after each click in order to get the food at the end of the session? How does it know what to do when the clicker clicks?

PY: Thus, the food rewards are not directly used in shaping the animal’s response, because the rate of food intake is not a controllable parameter.

RM: So the food has nothing to do with the training? I don’t get it.

PY: The clicker training method allowed the process of training to have nothing to do with deprivation and subsequent rewarding. Animals trained using a clicker were healthier, learned tricks more easily, and could even be made to improvise on cue.

RM: I would really like to know how this method works. And it would be great if you could give us an analysis of what’s going on from a control theory perspective.

PY: I think the take-home message of the book was that the clicker was key for the animal to understand how to behave, and why to behave that way. The other important thing was to focus on establishing an ideal training environment, in which there are

  1. no consequences to fear (no chance of punishment/negative-reinforcement)
  2. very simple, highly predictable parameters (in the form of the clicks) for the animal to seek from the trainer.

PY: I believe there is much to learn from this purified branch of behaviorism

RM: I agree. I think it would be really nice to have an anlysis of this training technique from a control theory perspective.

PY: , and I suggest that operant training (rather than conditioning) be understood as a very positive thing. There should be room for an analysis of a negative-feedback-only system in a positive-feedback-only environment. Furthermore, I think it’s interesting if PCT incorporated this theory of animal training,

RM: What do you think should be incorporated into PCT? I’d be surprised if PCT doesn’t already explain how and why her training system works.

PY: because it takes the emphasis off the PHYSICAL disturbance as the cause of the behavior (learning).

RM: I think if you provided a PCT analysis of the training technique it would help me understand what you mean when you say that the technique takes the emphasis off the disturbance as a cause of behavior. Disturbing a controlled variable is not the only way to control behavior. And in the case of learning new behavior, as is the case in animal training, disturbance resistance can’t really work at all since you want the animal to develop a new control organization and disturbance to a controlled variable only works with existing control organizations.

PY: This puts the emphasis on the nature of the reinforcement rather than on the nature of the disturbance, as far as learning is concerned. And I think learning should be more associated with a process of reinforcement, rather than as a response to a disturbance.

RM: Actually, using reinforcement as the basis of training is somewhat more coercive than using disturbance resistance. In order to use disturbance resistance to control behavior the animal already has to have very good control of the variable that is going to be disturbed; indeed, the better the animal can control a variable, the better you can control the actions the animal uses to protect the variable from your disturbances. Using reinforcement requires that the trainer become the only source of the reinforcement so that the reinforcement is provided only when the animal does the desired (by the trainer) behavior. But maybe Karen does something different than what is done in operant learning. If so, I’d be interested in what it is.

Best

Rick

···

Richard S. Marken, Ph.D.
Author of Doing Research on Purpose.

Now available from Amazon or Barnes & Noble

No virus found in this message.
Checked by AVG - www.avg.com
Version: 2015.0.5863 / Virus Database: 4328/9484 - Release Date: 04/08/15


No virus found in this message.
Checked by AVG - www.avg.com
Version: 2015.0.5863 / Virus Database: 4321/9479 - Release Date: 04/07/15

[From Rick Marken (2015.04.11.0915)]

···

 Bruce Abbott (2015.04.08.1030 EST)–

Â

BA: Philip Yeranostan’s description of clicker training is a bit off the mark. Clicker training uses the clicker to reinforce the behavior that the click followed. The clicker is known as a “bridging stimulus,� because it bridges the time between that act and the delivery of the primary reinforcer, allowing the animal to associate that reinforcer with the successful act.

 RM: Very helpful post, Bruce. Thanks.Â

 BA: I leave it as an exercise for the reader to explain these results in PCT terms.

RM: I’m waiting to see if any others will take you up on this; I think it’s an excellent suggestion for an exercise. If no one volunteers one in a couple days I’ll give it a try myself.Â

BestÂ

Rick

Richard S. Marken, Ph.D.
Author of  Doing Research on Purpose
Now available from Amazon or Barnes & Noble

[From Rick Marken (2015.04.15.0915)]Â

 Bruce Abbott (2015.04.08.1030 EST)--

Â

BA: Philip Yeranostan’s description of clicker training is a bit off the mark. Clicker training uses the clicker to reinforce the behavior that the click followed. The clicker is known as a “bridging stimulus,� because it bridges the time between that act and the delivery of the primary reinforcer, allowing the animal to associate that reinforcer with the successful act.

 RM: Very helpful post, Bruce. Thanks.Â

 BA: I leave it as an exercise for the reader to explain these results in PCT terms.

RM: I'm waiting to see if any others will take you up on this; I think it's an excellent suggestion for an exercise. If no one volunteers one in a couple days I'll give it a try myself. Â

RM: Well, it looks like there's not a lot of interest in this and I don't really have the time to give much of an explanation. But since you're the expert in this why don't you give the PCT explanation of what's going on.Â
BestÂ
Rick

···

--
Richard S. Marken, Ph.D.
Author of  <Amazon.com Research on Purpose.Â
Now available from Amazon or Barnes & Noble

[From Bruce Abbott (970816.0900 EST)]

I can't resist this one.

[From Bill Powers (970815.1129 MDT) --

Bruce Abbott (970815.1115 EST)

I think we're _close_ to getting at the heart of the disagreement, but this
isn't it. I know that food pellets don't appear until the animal presses
the lever. That is the direct, physical contingency. But the animal
presses the lever _because_ pressing the lever produces food pellets, and
_because_, in the past, the animal has learned that pressing the lever
_does_ produce food pellets.

Now this is a different matter: it is the existence of a _relationship_
that allows the animal to control the delivery of food pellets, and not
either component of the relationship alone. If there is anything that could
be called a reinforcer, it would be the relationship, an abstraction -- not
the food pellets. What we have here is animals detecting an abstract
relationship, and altering their behavior to take advantage of it.

I'm not sure why you want to call it "abstract." First you say it is a
relationship, then you say that a relationship is an abstraction, and
somehow that leads to the idea that we have an abstract relationship. But
if relationships _are_ abstractions, isn't calling a relationship "abstract"
redundant? It's a relationship.

You will answer that the animal presses the lever because it _wants_ the
food. That is true, too. But the animal wouldn't press the lever for this
reason unless lever-pressing produced the food and the animal knew that. So
there are many "becauses" to consider. The PCT analysis focuses on the
unobserved internal "because," whereas the Skinnerian analysis focuses on
the observable external "becauses," the latter because Skinner believed that
appeal to unobservable internal states had led psychology down the garden
path.

I give Skinner credit for sticking to observables, but not for his _a
priori_ assumption that the environment had to be the determinant of
behavior. That was what led his followers down a different garden path. The
basic problem is that there is no one "cause" of behavior or behavior
change. The organism obviously has to take the properties of the
environment into account, so the environment has an influence. But it is
the inner organization of the organism that determines what part the
environment will play in behavior.

Obviously, but Skinner wanted to avoid speculating about the inner
organization of the organism, because this type of speculation had already
established a very bad track record in psychology. One could always invent
just the right inner organization to explain the data at hand -- and people
were expending a lot of effort devising and carrying out tests to
discriminate different proposals, often with ambiguous results. Better to
carry out a systematic series of well-controlled empirical investigations --
the empirical relationships thus discovered would have lasting value; they
would be the phenomena that any competent theory of inner organization would
have to explain.

In reinforcement theory, the reinforcement phenomenon is simply a given, an
empirical fact. No attempt is made to explain why, under given conditions,
a given consequence of behavior will sustain that behavior.

And that, my friend, is the old red flag again. The consequence is not what
sustains behavior. Listen to your own advice. It is _the existence of the
relationship_ , plus the organism's continued need for the consequence,
that sustains the behavior. Every time you hint that somehow the food
pellets are doing something to the behavior, like "sustaining" it, I'm
going to blow the whistle. Right in your ear. Maybe this will condition you
to stop saying such things.

Maybe so, but I thought that by making clear what I _meant_ when I say such
things, you would grant me the license to use this more compact way of
speaking, rather than insisting that I have drifted back to meaning what
_you_ insist that I mean. When I speak of a given consequence of behavior
sustaining that behavior, I'm talking about the role of the contingency in
keeping the behavior going. I am not asserting that the consequence exerts
some magical force to propel the behavior. In a gasoline engine, the
gasoline sustains the engine's running. But so does the spark, and the air
being pulled into the cylinders, and the energy stored in the flywheel.
Remove any one of those things and the engine quits. When I speak of
"sustaining" behavior, I'm talking about the conditions necessary in order
to maintain its performance. Yes, there other necessary conditions besides
the provision of some particular consequence of the behavior; these are
hidden in the clause "under given conditions." But all I'm trying to convey
here is that under those conditions, that consequence is necessary: without
it performance will collapse. It would be nice if we could discuss the
issue with this understanding rather than insisting that I mean something
else and arguing on that basis.

What PCT does is propose a mechanism that explains why some events, when
made contingent on a given behavior, will act as reinforcers (or punishers).

Why are you doing this? Are you trying to drive me into Rick Marken's
position, of throwing up my hands and saying you just don't get it? PCT
does not explain any such thing! You're asserting that there are events
which, when made contingent on a given behavior, act as a reinforcer or a
punisher of that behavior.

Let's try to take this a step at a time. I'll take the prototypical example
of a rat pressing a lever for food pellets. The rat has been deprived of
food for a time and has previously learned that it can obtain food by
pressing a lever, whereby the apparatus delivers a food pellet, and then
consuming the pellet.

We start this particular session with the contingency in effect (presses
produce food.) We observe that under these conditions (food deprivation
plus contingency), the behavior (lever pressing-food consumption) is
maintained, i.e., we observe the rat repeatedly press the lever and consume
the food, press the lever and consume the food, press the lever and consume
the food.

Now break the contingency by disconnecting the feeder, so that pressing the
lever no longer produces a food pellet. After a time, we observe that the
rat is no longer pressing the lever at the high rate that had formerly
characterized its performance. In fact, it rarely presses the lever at all.

Now we establish the contingency again, and in short order we find the rat
pressing and eating, pressing and eating, as before.

_By definition_, the food pellet is a reinforcer. When the pellet is
produced by pressing the lever, pressing the lever occurs at a high rate.
When the pellet is not produced by the lever, pressing the lever is not
maintained. We might call this EAB's Test for the reinforcer.

Can PCT explain why the food pellet acts as a reinforcer? Bill Powers says
no. I say yes. The food deprivation prevents the rat from controlling its
daily food intake, which must be maintained because the food is the rat's
fuel and is being used up by metabolic processes. As a result, a controlled
variable (energy intake) is brought below its reference level. The rat has
learned that it can control the CV in the operant chamber by pressing the
lever in order to produce a food pellet and then consuming the pellet. The
pellet is able to increase the level of the CV (once the pellet is
swallowed) because it contains the necessary energy in the form of
digestable nutrients.

In the initial condition, the feeder produces a pellet when the lever is
pressed, closing the loop through the environment and thus establishing
control over the CV. The error in the CV (resulting from the previous
deprivation) now drives lever-pressing (and food consumption) and thereby
reducing the error in the CV. We observe sustained lever pressing.

Now we disconnect the feeder, thus opening the loop through the environment.
After a time, we observe that lever pressing has essentially ceased. Lever
pressing no longer allows the rat to control its CV, and some higher level
in the rat's system has stepped in to prevent the runaway lever-pressing
that would be expected of a simple one-level control system when its loop
has been opened on the environment side. So, in the steady state, we get
lever pressing when lever pressing produces food, and we do not get lever
pressing when lever pressing does not produce food, the pair of outcomes
which together define the food as a reinforcer. We have explained why the
food is required in order to sustain lever pressing (is that better, Bill?
I almost wrote "why the food sustains lever pressing.") And we have
explained why, if we are to observe this outcome, it is necessary to deprive
the animal of food. And you say that PCT doesn't explain the reinforcement
phenomenon? I think I've just shown otherwise.

I hope I haven't driven you into Rick's camp. I understand the mosquitos
are pretty bad over there . . . (;->

Regards,

Bruce

[From Bill Powers (970816.1007 MDT)]

Bruce Abbott (970816.0900 EST)--

I give Skinner credit for sticking to observables, but not for his _a
priori_ assumption that the environment had to be the determinant of
behavior...

Obviously, but Skinner wanted to avoid speculating about the inner
organization of the organism, because this type of speculation had already
established a very bad track record in psychology. One could always invent
just the right inner organization to explain the data at hand ...

Skinner was quite right in criticizing the "intervening variable" theories,
those which explained behaviors using Bateson's "Dormitive Principle"
(sleeping powders make you sleepy because they contain a Dormitive
Principle). But that kind of theory is not a system analysis or a system
model. What Skinner should have been criticizing was the lack of
understanding in psychology of the way more successful sciences make
models. But he couldn't do that, because he didn't know how to make models,
either.

-- and people
were expending a lot of effort devising and carrying out tests to
discriminate different proposals, often with ambiguous results. Better to
carry out a systematic series of well-controlled empirical investigations --
the empirical relationships thus discovered would have lasting value; they
would be the phenomena that any competent theory of inner organization would
have to explain.

You are really idealizing Skinner's practices, or else describing what he
advocated rather than what he did. I admire his cleverness, but not his
ability to think up critical experiments. He was much too ready to
compartmentalize his thinking, so that for example he could demonstrate how
reducing the reinforcement rate could drive an animal to greater and
greater behavior rates -- without ever noticing that in this case,
_reducing_ reinforcement was _increasing_ behavior. Skinner never seemed to
understand that a general principle is supposed to be _general_ -- apply
all the time, not just when you want it to or when it happens to fit what
you see. If an increment of reinforcement is supposed to produce an
increment in behavior, it must do this EVERY TIME, even when you're not
focusing on that aspect of the experiment. If it doesn't, it's not a
general principle.

In reinforcement theory, the reinforcement phenomenon is simply a given, an
empirical fact. No attempt is made to explain why, under given conditions,
a given consequence of behavior will sustain that behavior.

And that, my friend, is the old red flag again. The consequence is not what
sustains behavior. Listen to your own advice. It is _the existence of the
relationship_ , plus the organism's continued need for the consequence,
that sustains the behavior. Every time you hint that somehow the food
pellets are doing something to the behavior, like "sustaining" it, I'm
going to blow the whistle. Right in your ear. Maybe this will condition you
to stop saying such things.

Maybe so, but I thought that by making clear what I _meant_ when I say such
things, you would grant me the license to use this more compact way of
speaking, rather than insisting that I have drifted back to meaning what
_you_ insist that I mean.

What I want to do, Bruce, is deprive you of the use of language that
permits you to assert environmental causation, whether you intend this to
be the meaning or not. Since I can't do that, I am trying to persuade you
to give it up voluntarily, one day at a time. This isn't just a "compact
way of speaking" -- it's a way of getting two messages across at the same
time, one overt and one covert. The overt meaning, the one you will admit
to, is the whole situation in which providing a contingency results in an
animal using it for control of something it wants. The covert meaning is
that the thing it controls is really controlling _it_. All you have to do
is to read the literature of EAB to know that this covert message is
explicitly intended by most workers. Why encourage this view? Why create
the appearance that you support it?

When I speak of a given consequence of behavior
sustaining that behavior, I'm talking about the role of the contingency in
keeping the behavior going. I am not asserting that the consequence exerts
some magical force to propel the behavior.

Then choose your words carefully so you don't give the impression that some
variable involved in behavior is keeping the behavior going. What keeps the
variable going? The behavior! Skinner wanted to treat a dependent variable
as if it were an independent variable; you can't get away with that among
engineers, or in many other quarters.

In the first sentence above you say first that a _consequence_ of behavior
sustains the behavior, and immediately afterward, in the same sentence,
that it is the _contingency_ that keeps the behavior going. The contingency
is not a consequence of behavior; it's a property of the environment lying
between the behavior and its consequences. You're using these terms in a
loose free-associative way. If you mean contingency, which is a function
making one variable depend on another, say contingency. If you mean
consequence, which is a variable that is a function of another variable,
say consequence. The clarity of your thinking is reflected in the clarity
of your language. If you use muddled language, your thinking is muddled,
too. And even if you manage to keep clear thinking going despite the
muddled language, your communication gets muddled whether you know it or not.

In a gasoline engine, the
gasoline sustains the engine's running. But so does the spark, and the air
being pulled into the cylinders, and the energy stored in the flywheel.
Remove any one of those things and the engine quits. When I speak of
"sustaining" behavior, I'm talking about the conditions necessary in order
to maintain its performance.

There is a difference between necessary and sufficient. It is necessary for
behavior to be able to affect consequences if there is to be any change in
behavior, but it is not _sufficient_ for behavior to affect consequences.
Many other variables and functions of equal importance are involved, no one
of which is sufficient, but each one of which is necessary. We are not
talking about lineal causation here; we're talking about a circle of
causation, a closed loop in which all the parts play a role in determining
the outcome, but no one part plays a determining role. A consequence of
behavior is merely one link in a closed loop of variables which are
successive functions of previous variables in the loop, all the way around
it. Ultimately, a consequence of behavior, if controlled, is a function of
itself. Every variable in the loop is a function of itself. The only
independent variables are disturbances and reference signals.

Yes, there other necessary conditions besides
the provision of some particular consequence of the behavior;

Again, muddled language. Are you talking about the _provision_ that makes
the consequence a function of behavior (the contingency)? Or are you
talking about the _consequence_ that is provided by the behavior? WHAT
provides the particular consequence of the behavior, and what do you mean
by "particular?" It is the behavior that provides the particular _amount_
of the consequence; it is the contingency that provides a particular _kind_
of consequence. The experimenter setting up the contingency has no way to
determine how much of a consequence will appear; only what kind. It is the
organism's behavior that, given the contingency, determines how much (if
any) of the consequence will appear. By leaving out the provider of the
consequence, you make it appear that the consequence just appears, like an
independent variable. But for the phenomenon you're talking about to exist,
it is _essential_ that it be the behavior that provides the consequence,
not some independent unnamed agency.

This subjectless construction is a way of concealing the fact that it is
the organism that produces a particular amount of consequence. To
acknowledge that cause of the consequence explicitly would weaken the idea
that the consequence has some causal role in shaping behavior. Therefore,
the propagandist hides the agent by skipping over it in the sentence, as if
it made no difference what produced the reinforcer.

It would be nice if we could discuss the
issue with this understanding rather than insisting that I mean something
else and arguing on that basis.

I refuse. Say what you mean, explicitly and precisely. Why should I always
have to be clarifying and translating your language for myself? If you
insist on continuing to use your "compact" language, it can only be because
you _want_ to get the covert message across -- you want to leave an opening
for an interpretation that puts causation into the environment. You want to
sound like a behaviorist. Come on, Bruce; what's wrong with clarity and
precision?

···

-------------------------

Let's try to take this a step at a time. I'll take the prototypical example
of a rat pressing a lever for food pellets. The rat has been deprived of
food for a time and has previously learned that it can obtain food by
pressing a lever, whereby the apparatus delivers a food pellet, and then
consuming the pellet.

We start this particular session with the contingency in effect (presses
produce food.) We observe that under these conditions (food deprivation
plus contingency), the behavior (lever pressing-food consumption) is
maintained, i.e., we observe the rat repeatedly press the lever and consume
the food, press the lever and consume the food, press the lever and consume
the food.

To say that the behavior "is maintained" is not the same as explicitly
saying that the _rat_ maintains its behavior, and thus maintains the food
intake. This passive construction leaves open the way to speculate about
WHAT is doing the maintaining. But the maintenance is not done by anything
outside the loop: the whole process _maintains itself_.

Now break the contingency by disconnecting the feeder, so that pressing the
lever no longer produces a food pellet. After a time, we observe that the
rat is no longer pressing the lever at the high rate that had formerly
characterized its performance. In fact, it rarely presses the lever at all.

Now we establish the contingency again, and in short order we find the rat
pressing and eating, pressing and eating, as before.

_By definition_, the food pellet is a reinforcer. When the pellet is
produced by pressing the lever, pressing the lever occurs at a high rate.
When the pellet is not produced by the lever, pressing the lever is not
maintained. We might call this EAB's Test for the reinforcer.

I'm beginning to think that your problem is simply one of ambiguous
thinking, and it shows in your language. Your concepts are fuzzy. "When the
pellet is produced by pressing the lever" could mean either "when a pellet
appears because of a lever press" (emphasis on the appearance of a pellet)
or "when it is possible that a lever press might produce a pellet"
(emphasis on the existence of the contingency). Likewise, "When the pellet
is not produced by the lever" could mean "When pressing the lever produces
no pellet," or "when there is no press that produces a pellet," or "when it
is not possible that pressing the lever could produce a pellet." Your words
and constructions do not pin you down to a single meaning, which indicates
to me that the underlying concepts are not clearly distinguished in your
mind. I don't think you have clearly distinguished, in your own thinking,
between the appearance of a food pellet (the reinforcer) and the
relationship that gives behavior the ability to produce a food pellet (the
contingency). You don't seem able to make up your mind whether it is the
occurrance of the reinforcer that accounts for the increased behavior rate,
or the existence of the contingency.

You could equally well interrupt the loop at any other point and define any
other variable as the reinforcer. For example, you could inject curare into
the muscles to prevent bar-pressing. No muscle contractions, no behavior;
muscle contraction, behavior -- but only if the loop is closed. So the
muscle contractions are reinforcers. Or you could cut off sensory access to
the food pellets; same result: sensory responses are reinforcers -- if the
loop is closed. Heck, you could clamp some variable in the middle of the
program that generates the contingency to zero, and show that that variable
is a reinforcer. The whole loop is made of nothing but reinforcers!

Can PCT explain why the food pellet acts as a reinforcer? Bill Powers says
no. I say yes. The food deprivation prevents the rat from controlling its
daily food intake, which must be maintained because the food is the rat's
fuel and is being used up by metabolic processes. As a result, a controlled
variable (energy intake) is brought below its reference level. The rat has
learned that it can control the CV in the operant chamber by pressing the
lever in order to produce a food pellet and then consuming the pellet. The
pellet is able to increase the level of the CV (once the pellet is
swallowed) because it contains the necessary energy in the form of
digestable nutrients.

The food intake is not maintained "because the food is the rat's fuel." It
is maintained because there is a physical reference signal specifying a
nonzero level of food intake (some of the time) for the associated control
system (and, of course, because the rest of the control system is working).
And that reference signal is not maintained "because the food is the rat's
fuel." It is maintained because some other system is controlling stored or
circulating glucose and is experiencing an error. And that system does not
maintain the glucose level "because the food is the rat's fuel." It
maintains it because it receives a nonzero reference signal from still
another system perceiving and controlling effects that depend on blood
glucose level. And so forth. However far back you follow this trail, you
will never come to a control system that is acting "because the food is the
rat's fuel." That sort of idea is irrelevant to understanding how the rat
works, what it is controlling, or why it does what it does.

So, in the steady state, we get
lever pressing when lever pressing produces food, and we do not get lever
pressing when lever pressing does not produce food, the pair of outcomes
which together define the food as a reinforcer. We have explained why the
food is required in order to sustain lever pressing (is that better, Bill?
I almost wrote "why the food sustains lever pressing.")

See above discussion of what is a reinforcer. And no, it is not better. You
still leave out what is sustaining the lever pressing. What sustains the
lever pressing is a signal from the nervous system that depends both on
food intake and the reference signal setting. WHAT requires food "in order
to sustain lever pressing"? This damnable passive construction in which the
subject is quietly dropped out! It is the _RAT_, via reference signals,
that "requires" the food.

And we have
explained why, if we are to observe this outcome, it is necessary to deprive
the animal of food. And you say that PCT doesn't explain the reinforcement
phenomenon? I think I've just shown otherwise.

PCT explains the behavioral relationships, but it doesn't explain
reinforcement because there is no such thing. The food has no special
reinforcing effect; it has a nutritional effect, and that is all. Any
reinforcing effect is a product of the observer's imagination. I refuse to
attribute any effect to the food other than its effect as a source of
energy and building materials. Those are the ONLY effects we need to
consider to apply the PCT model.

I agree that the manipulations you describe above can easily give the
impression that the food is doing something in addition to providing
nutrition. That, indeed, is the whole thrust of your argument. If you let
behavior produce food, you get behavior; if you don't, you don't get
behavior. What could be more reasonable than to conclude that the food is
what determines the presence or absence of behavior? But this naive
conclusion focusses down on just one variable in the loop, and overlooks
the fact that you could say exactly the same thing about ANY variable. The
role you are calling "reinforcer" could be assigned to any variable in of
the loop, as could the role you call a "contingency" (to any function in
the loop).

What you're doing is known as synecdoche -- "a figure of speech in which a
part is put for the whole, as in 50 sail meaning 50 ships." In this case,
you are giving a part a function which is a property of the whole loop, a
part that you happen to be able to see and consider important.

I don't know about this discussion, Bruce. I find myself giving you lessons
in how to speak and think, lessons to which I very much doubt that you are
paying any attention. You have your way of talking about things, and it
doesn't seem that anything can influence that or lead you to wonder if
these ideas are really as clear as you believe they are.

Best,

Bill P.

[From Rick Marken (951201.1930)]

Bill Powers(951201.1430 MST) to Bruce Abbott --

If you introduce disturbances in an operant conditioning experiment, you
will no longer be able to define a reinforcer, because no state of the
variable in question will increase the probability of any particular
response. In fact, the responses can be made independent of the state of
the reinforcer, yet control will continue.

I have repeatedly brought this point up, without a comment from you. If
you will promise to discuss this point, I will work up a simple experiment
to illustrate it. But if you're just going to let it slide past again,
I won't bother. What about it?

Bruce Abbott (951201.2105 EST)--

Ah, you cut me to the quick! I have indeed commented: check your archive.
I don't remember which post, but my reply was given the last time you
brought up the topic. As I recall, you asked me about it in the form of a
test. After reading my reply, you said I passed.

I think Bill is making a VERY important point, here; what he is saying
is that when you introduce disturbances to a variable (like amount of
food) that seems to be a reinforcer (one that seems to increase the
probability of a response) you will see that the variable doesn't
actually work as a reinforcer because no state of the variable "will
increase the probability of any particular response". Bill is saying,
in other words, that reinforcement (the apparent strengthening influence
of certain consequences on the reponses that produce them) is an illusion.

I would like to hear your reply to Bill's point (even if you did answer
before; I don't remember it and it's pretty tough to look through the
archives for this kind of stuff). In particular, I would like to know
whether you think it would be worth it for Bill to "work up a simple
experiment" to illustrate his point. In other words, is it worth it to
try (again) to demonstrate that consequences don't reinforce?

Best

Rick

[From Bruce Abbott (951203.1200 EST)]

Rick Marken (951201.1930) --

Bill Powers(951201.1430 MST) to Bruce Abbott --

If you introduce disturbances in an operant conditioning experiment, you
will no longer be able to define a reinforcer, because no state of the
variable in question will increase the probability of any particular
response. In fact, the responses can be made independent of the state of
the reinforcer, yet control will continue.

I have repeatedly brought this point up, without a comment from you. If
you will promise to discuss this point, I will work up a simple experiment
to illustrate it. But if you're just going to let it slide past again,
I won't bother. What about it?

Bruce Abbott (951201.2105 EST)--

Ah, you cut me to the quick! I have indeed commented: check your archive.
I don't remember which post, but my reply was given the last time you
brought up the topic. As I recall, you asked me about it in the form of a
test. After reading my reply, you said I passed.

RM:

I would like to hear your reply to Bill's point (even if you did answer
before; I don't remember it and it's pretty tough to look through the
archives for this kind of stuff). In particular, I would like to know
whether you think it would be worth it for Bill to "work up a simple
experiment" to illustrate his point. In other words, is it worth it to
try (again) to demonstrate that consequences don't reinforce?

If you insist. I can understand how everyone has forgotten our previous
exchange on this issue; after all, it's been a whole MONTH since we
discussed it. Rather than repeat myself, I append my previous reply below.
Bill Powers's assessment follows it.

Bruce

···

-----------------------------------------------------------------------------
[From Bruce Abbott (951028.1110 EST)]

Bill Powers (951028.0600 MDT) --

Bruce Abbott (951027.1600 EST) --

You seem to have been reading ahead in the text, and have somehow
managed to complete PCT 301, which is pretty close to all there is.
However, there is still one outstanding paper from 101, concerning
reinforced behavior in the presence of unsystematic disturbances of the
reinforcement rate.

Dear Professor Powers:

Enclosed is my final paper for PCT 101. I hope you can get it graded soon
so I know what courses to take next semester.

Sincerely,

Bruce

----------------------------------------------------------------------------
  REINFORCED BEHAVIOR IN THE PRESENCE OF DISTURBANCE TO REINFORCEMENT RATE

                            Bruce Abbott
                              PCT 101

In my previous paper I described how reinforcers can be defined as events
that, when made contingent on an action, tend to reduce error in the system
producing that action. In the prototypical operant procedure, error is
produced by food deprivation and the event that tends to correct this error
is delivery of food. If the error (deprivation level) is constant the
output function will develop a particular level of output (e.g., rate of
lever-pressing). That rate of output will produce a given rate of food
delivery, depending on the schedule imposed by the experimenter.

If the rate of food delivery is itself a controlled perception, then
disturbances to this rate will be opposed. One way to disturb food rate is
to interpose free (noncontingent) food deliveries between those produced by
the rat's lever-pressing actions. The extra deliveries would be expected to
raise the rate of food delivery above its reference level, developing an
error in the food-rate control system (perceived rate above reference rate).
This would produce a reduction in the rat's rate of lever-pressing and thus,
through the schedule parameters, a reduction in food-rate, reducing the
error. Withholding some proportion of scheduled (response-contingent) food
deliveries would be expected to have the opposite effect, reducing food rate
below reference and producing a compensatory increase in response rate. In
consequence the overall rate of food delivery will remain roughly constant
despite these experimenter-imposed distrubances.

The expected result is paradoxical from the point of view of reinforcement
theory, if one assumes that the disturbances have little effect on
deprivation level: interspersed free food deliveries would appear to
suppress responding on the lever, whereas deleted food deliveries would
appear to reinforce responding.

In order to simplify the explanation, I have ignored several important
details, such as the fact that the observed performance (lever pressing)
would involve several levels of control; in addition we have evidence that,
in at least one situation (ratio schedules), one kind of disturbance to
reinforcement rate (changing the ratio requirement) did not appear to
produce the expected compensatory changes in response rate (although we need
better data to confirm this conclusion). The predictions offered here are
therefore intended only to indicate how control theory can be applied to the
problem of disturbance to reinforcement rate, under the assumption that
reinforcement rate is in fact a controlled perception under these
conditions. Developing a proper model of actual rat behavior under these
schedules will require additional research to determine what perceptual
variables are in fact being controlled and how the various control systems
involved interact.
------------------------------------------------------------------------------
[From Bill Powers (951028.1150 MDT)]

Bruce Abbott (951028.1110 EST) --

A+.

The "A" is for stating the theoretical prediction correctly, and the "+"
is for noting correctly the research that remains to be done.
------------------------------------------------------------------------------