[From Bruce Abbott (950702.1520 EST)]
I think we're making progress.
Bill Powers (950702.0500 MDT) --
Bruce Abbott (950701.1355 EST)
What we're concerned with, then, is the "motivational" aspect of
reinforcement, which covers the right (descending) side of the
behavior/reinforcement curve.
Why just the right side? The motivational aspect should cover the whole
curve, since we're dealing with steady-state behavior.
The second role is as the source of motivation for executing the
behavior which has been learned. The behavior may have been
"selected" (first role), but if there is no motivation it will not
be executed. In the steady state (when "selection" has been
completed), the rate of behavior will vary with the level of
motivation.
This is unclear. If the reinforcement is the source of motivation, then
how does motivation vary with reinforcement rate? Does motivation
increase as a function of increasing reinforcement rate?
If we consider food pellets as reinforcers, then for a given level of
deprivation (held constant!) and a given type and size of pellet, etc., the
motivation attributed to the reinforcer would be constant across rates.
However, responding brings costs which may vary as a function of response
rate. These costs are effectively a source of motivation for not
responding. The observed rate of responding would reflect a sum of these
opposing motivational levels.
Or is
motivation an independent variable determined in some other way?
See below.
Apparently, a reinforcer can be present in the role of "selecting" a
given behavior, without at the same time automatically "motivating" it.
No, to provide selection, the reinforcer must also motivate. However, once
the steady state is reached, there can be motivation without further selection.
During steady-state performance, "selection" is complete, so the
only influence on performance is the level of motivation.
What do you mean by "performance?" Behavior rate?
Performance is behavioral output, which would cover all facets, including
such aspects as rate, duration, response force, even temporal pattern.
If reinforcement is
the source of motivation, then can't we just say that behavior rate goes
with reinforcement rate?
No. See above.
Does the term "motivation" have a meaning other
than rate of reinforcement? Is it possible to have high motivation and
low reinforcement rate, or high reinforcement rate and low motivation?
Or is motivation merely, to a first approximation, proportional to
reinforcement rate?
One could have high motivation (say, because of high deprivation) and low
reinforcement rate (because of schedule parameters). If the reinforcer can
be earned with low effort and there are no alternative sources of
reinforcement, fairly high rates might be sustained at relatively low levels
of motivation.
Satiation is not a rate but a state, and if you have reached
satiation, there is no motivation for behavior whose consequence is
to produce food.
Satiation, then, works to reduce motivation: i.e., the closer the
organism approaches to satiation, the less motivation there is. So
motivation depends on a state of the organism that is affected by the
reinforcer but that can vary independently of the reinforcer.
Yes.
The rate of lever-pressing observed on the VR schedule will thus
represent a dynamic equilibrium involving the reinforcing value of
the reward, the response cost, the value of competing sources of
reinforcement, and the degree to which lever-pressing and the other
behaviors motivated by those competing sources of reinforcement
interfere with each other.
If an organism is receiving reinforcement from one source, how does the
source not providing reinforcement at the moment "compete" with the
source actually providing the reinforcement? If only one behavior at a
time can be performed, how does the one not being performed "interfere"
with the one actually being performed?. If reinforcement is not being
received from a given source, and the behavior associated with that
source is not being performed, is the behavior not being performed still
being "motivated" by the reinforcer not being received?
For lever-pressing to be motivated by a food pellet, the rat must have had
some experience through which the act of lever pressing has become
associated with the consequence of receiving the pellet. Because of this
association (learning), the food pellet motivates lever pressing and not
some other behavior (selective property). The motivation to press the lever
is thus "there" throughout the session, as are any other motivations to do
other things. It is not necessary for the reinforcement from a given source
be "received" continuously in order for that reinforcer to motivate the
behavior with which it has become associated throught prior experience.
If only one behavior at a time can be performed, then two possibilities
present themselves: (a) the behavior with the highest "payoff" may be
performed to the exclusion of the others (exclusive choice); (b) the
behaviors may be alternated in some fashion so that reinforcement from
different sources can be "collected." In the case of (b) it may be possible
for the behaviors to alternate in such a way that that essentially all
opportunities for reinforcement from each source can be taken advantage of.
Alternatively, there may be constraints such that collecting reinforcers
from one source causes missed opportunities for collecting them from the
other source. In the latter case there is competition between reinforcement
sources in the sense that increasing the rate of reward from one source
decreases the rate of reward avaliable from the others. What happens in
these different cases has been the subject of considerable research.
Now, what determines the reinforcing value of the food pellets,
that is, the effectiveness of the reinforcer in maintaining the
behavior that produces it? Again the situation is somewhat
complex. All else being equal, the rat will respond more
vigorously for a pellet of food the higher the level of deprivation
(up to the point where deprivation begins to weaken the animal, a
region rarely explored in operant studies). So higher deprivation
leads to the pellet being more highly valued, capable of sustaining
a higher response rate.
Deprivation, then, multiplies the response-producing power of each
reinforcement by some factor. It is not, however, reinforcing in itself,
because if behavior is not maintained by reinforcement, deprivation
would not be able to affect behavior. On the other hand, it would also
follow that reinforcement by itself is not effective unless there is
deprivation, because of the multiplicative relationship.
Yes.
The size of the pellet has a similar effect (note: both effects are
nonlinear. Each successive doubling of pellet size does not double
the value of the pellet; a log function may be indicated.) The
flavor of the pellet would also be a factor determining value, and
several other things as well that I will skip over in the interest
of brevity.
So physical and biochemical characteristics of the pellet have an effect
similar to that of deprivation. Is this also a multiplicative effect?
Yes and no. Size yes (using perhaps log size rather than absolute size).
Taste enters in, and it is possible to get motivation without any nutritive
value (as in saccharine), so all factors are not strictly multiplicative.
In addition, there is contiguity between the response and the
delivery of the pellet to consider when determining value. The
larger the delay between responding and reinforcement, the weaker
the effect of the reinforcer on the response (the so-called
gradient of reinforcement).
Is contiguity simply the reciprocal of delay?
Yes.
How is this delay different from a ratio schedule?
One could program a ratio schedule in which there was a fixed delay between
the completion of the ratio and the delivery of the food pellet. Longer
delays will produce lower rates of responding on the schedule. But delay
may also enter in when considering the effect of the ratio requirement
itself. The first response in the ratio will occur at a delay to
reinforcement that depends on the number of additional responses that need
to be completed and the rate at which those responses are made. If that
rate (the running rate) were constant across ratios, then larger ratios
would impose longer delays between the first response in the ratio and
reinforcer delivery. During the postreinforcement pause (after the pellet
has been consumed) the animal has a choice: return to the lever or engage in
some other behavior (even if that behavior is just to rest). The
probability that the rat would return to the lever would depend on the
operant levels of these competing behaviors, which in turn would depend on
the relative value of immediate reward for the alternative behavior versus
the delayed reward for lever pressing. The prediction would be that larger
ratios should be associated with longer postreinforcement pauses.
If a response occurs, another response
that occurs within the delay time will go unrewarded, so there may be
several responses before the reinforcer occurs. The effect is, on the
average, the same as rewarding only one response out of n, where n is
the average number of responses made during the delay time.
That would depend on whether the rat discriminates the lack of contingency
between response and reinforcement during the delay.
As for the delay imposed by the ratio requirement itself, all responses are
required before reinforcement will be delivered, so this is a different case
from responding during an imposed delay. Once the rat has returned to the
lever, it is easier to continue lever-pressing than to break off and do
something else. Furthermore, once the first response has occurred, a
smaller ratio remains and thus a shorter delay to reinforcement. There is
thus a strong tendency to continue responding once responding on the lever
has begun. The usual view, I think, would be that each response is
reinforced by the eventual delivery of the food pellet, but at different
delays, so that the amount of reinforcement would be slightly greater for
each subsequent response in the series. This would predict that responding
should accelerate during completion of the ratio. This acceleration is
sometimes observed during acquisition, but disappears in the steady state.
A proposed explanation is that with sufficient experience the rat begins to
discriminate the relationship between _rate_ of responding and _rate_ of
reinforcement (as opposed to merely the occurrence of a response, and the
occurrence of reinforcement), and then begins to respond so as to maximize
the rate of reinforcement (but with response rate diminished by the effect
of response costs).
This holds
for interval as well as ratio schedules. Is this reinforcement gradient
used as an explanation for post-reinforcement pauses?
Yes, as described above.
If there were no other sources of reinforcement in the environment,
and if lever-pressing had no cost, one might expect the rat to
press the lever at maximum rate so as to maximize the rate of
reinforcement, regardless of the level of motivation.
So if we reduced the cost of responding nearly to zero, and eliminated
all alternative sources of reinforcement but the one kind of response,
the animal would be pressing at its maximum possible rate. This suggests
that on a FR1, where the cost per reinforcement is low, the animal would
be pressing faster than on a FR 10, where the cost per reinforcement is
high;
Yes, if that were the only consideration. The problem is that the lower
rates that the higher ratios would sustain would substantially reduce the
overall rate of reinforcement. If the rat were sensitive only to local
contingencies (dependence of individual reinforcer delivery on individual
responses), this would be the clear prediciton. But empirical tests
indicate that rats can learn to discriminate rates of reinforcement as well
as delays to individual reinforcers. Assuming that a higher rate of food
delivery is more reinforcing than a lower rate (to a hungry rat), a
prediction based on overall rates would be that rate would increase so as to
maximize reinforcement rate (but with rate of reinforcement reduced to
somewhat less than maximum in proportion to response cost, which would
reduce the effective value of a given rate of reinforcement).
I know what you're thinking--the theory tries to have it both ways: it can
predict either a direct or inverse relationship here. But the situation is
really no different from the one faced by a PCT analysis of the same
performance. What PCT model you build depends critically on which variable
or variables are being controlled. Logically, several are possible, and
different choices will produce models that behave differently. Whichever
theory you subscribe to, your efforts will be focused on identifying the
variables actually involved and determining how they contribute to the final
performance. And, of course, you will try to develop a model that is
consistent with the data. This criterion will allow you to rule out certain
alternative constructions of the model.
and that if the nutrient value of a pellet were increased, thus
requiring less behavior to produce the same value, the rate of pressing
would also increase.
Well, not the _nutrient_ value per se, but yes, increasing the size, for
example, would be expected to increase the rate, at least under a local
analysis, and assuming that satiation is not approached during the period of
observation.
By the time you get to an analysis based on overall rate of reinforcement
and rate of responding, it is but a short step to a regulatory view. Rate
of reinforcement might, for example, be taken to mean rate of intake, which
would factor in pellet size as well as rate of delivery. Close the loop,
add a set point, and viola, PCT.
The suggestion is also that if there were no (or little) cost and no
competing reinforcers, the organism would maximize the rate of
reinforcement regardless of the level of deprivation (which is the
motivation that gives reinforcers their effect). So even if replete, the
animal would still press as fast as possible.
Well, if the equation is a multiplicative one, at zero motivation, one has
zero output. But under these ideal (unreal) assumed conditions, even an
infinitessimal level of motivation would drive the output to maximum. But
what real system has infinite gain?
So far, here are the general ideas I get from this.
1. Reinforcers play two roles: as "selectors" of behavior and as
"maintainers" of behavior.
Yes.
2. Reinforcers, even when not actually being received, can "compete"
with reinforcers actually being received.
Yes.
3. Behaviors, even when not being performed, can "interfere" with other
behaviors actually being performed.
No. (see discussion earlier in this post) But the occurrence of one
behavior can interfere with (prevent) another behavior, so that it is not
performed, or is performed at a reduced rate compared to the rate if
interference were not present.
4. When reinforcement is maintaining a steady state of behavior,
"motivation" can still vary and thus vary the amount of behavior being
maintained by the reinforcement. So motivation and reinforcement can
vary independently.
Careful--"reinforcement" and "reinforcer" are not equivalent terms. The
reinforcer is not affected by deprivation (it's still the same food pellet).
But its motivating properties relative to the organism being subjected to
deprivation are.
5. Motivation depends negatively on a state of the organism called
"satiation". An approach to satiation reduces motivation.
Yes, if we're talking about some quantity associated with satiation effects.
Not all do: footshock is an example of one that doesn't.
6.Deprivation multiplies the ability of reinforcers to maintain behavior
by altering the "value" of a given reinforcer, as well as increasing
motivation.
Yes, but be careful here. Deprivation might affect the value of different
reinforcers to different degrees (different parameter-values relating
deprivation to value). The motivation to obtain any of these reinforcers
would increase with deprivation, but the motvation to obtain any particular
reinforcer IS its value.
7. Physical characteristics of the reinforcer have effects similar to
those of deprivation: increasing the size of a pellet will have an
effect similar to that of increasing deprivation, increasing the value
of the reinforcement and thus leading to more behavior.
Yes.
Note: the terms put in quotation marks the first time they are used are
undefined explanatory terms.
Well, perhaps I didn't define them in my post, but that is not the same as
their being undefined.
There is at least one puzzling relationship among these purported facts.
If, as you say, each one of them has been established through
experimental observation, this makes them all the more puzzling.
It has been observed that deprivation increases the value of a given
reinforcer, and increasing the value of a reinforcer is observed to
increase the level of behavior maintained by that reinforcer. But the
behavior that is maintained by reinforcement works in the direction of
reducing deprivation, which is obverved to reduce the value of the
reinforcer. At the same time, reducing deprivation decreases motivation
by causing an approach to satiation, which would also work against the
effect of the reinforcement.
Yes, and that is what is observed. However, operant studies in which
satiation might be a problem are designed to minimize this effect by keeping
reward size small and/or session length short.
Motivation is increased by deprivation, so
it is increased by a decrease in obtained reinforcement. But a decrease
in motivation would decrease the effectiveness of a reinforcer in
maintaining behavior, and that would tend to decrease reinforcement and
thus increase motivation.
You are considering only the case in which reinforcement rate is reduced
below that necessary to maintain the current level of deprivaiton. However,
in typical studies a decrease in obtained reinforcement would not increase
deprivation, it would slow the rate at which satiation occurred. The level
of motivation would not increase, it would decline more slowly.
There is a simple explanation of these somewhat contradictory
observations: what motivates behavior (at least after the selection
phase is over) is deprivation, not reinforcement. The more reinforcement
that is obtained, the less the deprivation becomes, and the less
behavior is observed. When the deprivation is completely removed,
behavior ceases. If reward size is increased, the deprivation is removed
faster and behavior ceases sooner. In the limit, a behavior that
produces the entire day's supply of reinforcer in one lump will suffice
to terminate behaving after one response.
A couple of posts ago I suggested what I thought would be the PCT
interpretation of what I had been describing in reinforcement terms. Your
response was to inform me that you knew what PCT would predict, as if I had
been attempting to instruct you on the principles of PCT. I was hoping you
would tell me whether you agreed with my analysis. I now see that you do.
Good!
Maintaining a general deprivation over a 24-hour period does not mean
that short-term deprivations cannot be created and removed. Removal of
short-term deprivation explains why even hungry rats normally eat in
bouts.
I agree.
This is, as you will recognize, the PCT explanation. This approach
focuses not on the range of reinforcements between zero and the obtained
amount, but on the range between the obtained amount and the satiation
level -- in other words, on the error (I am now accepting your
definition of satiation). The reference point is shifted from the
physical zero to the satiation level of obtained reinforcement.
Yes, but let's not forget that we're talking about at least two levels of
control here. The reference for _rate of reinforcement_ (what Motheral's
graphs show) will be zero when the reference for _nutrient level_
(satiation point) has been reached.
That's going to have to be it--this program won't let me enter more than
this in a single post. Anyway, that's probably enough for now.
Regards,
Bruce