Ratio Data

[From Bruce Abbott (950630.1550 EST)]

Bill Powers (950629.1810 MDT) --

    Bruce Abbott (950629.1530 EST)

    The "data for the first set" did not use running rate rather than
    the mean response rate, they used mean response rate. What you've
    done here is to compute the mean response rates based on the mean
    running rates and postreinforcement pause lengths. Here are the
    results from the two sources compared (averages only, for
    illustration):

Ok. Actually, you did pretty well estimating from graphs. The revised
graph for the second set is . . .

Bill, I think you understand, then you say something that seems to indicate
that you don't follow at all. The first and second set ARE THE SAME DATA.
Where they differ, it's just because of the uncertainty of my estimation
from the graphs. Therefore, when you say

Interesting that about the same reference level is implied.

you are saying essentially that it is interesting that the same data imply
the same reference level. It's not interesting at all.

Imagine that at each X you have the values of three variables, Y, A and B.
You have a graph showing each variable as a function of X. Reading from the
graph, you estimate Y, A, and B. But Y = A + B. Thus a plot of Y and a
plot of A + B should be identical except for errors in estimating the values
of Y, A, and B. It should come as no surprise at all that the graphs of Y
vs X and (A+B) vs X follow almost identical functions.

    Where behavior influences perception which influences behavior this
    notion of directional influence loses meaning, but I see nothing
    wrong with such statements as, "given a particular reference level,
    adding disturbance X to the controlled perceptual variable causes
    the system output Y to change."

I'll accept the term "influence" because I've always understood that to
mean a non-exclusive contribution to the final result. But to use
"cause" when more than one variable affects the outcome, it seems to me,
gives the wrong impression.

How about "effect"? Most researchers use "effect" as a synonym for "influence."

    On what basis do you say the "entire data set ... shows effects of
    satiation"?

Here is the relationship between reinforcements and behavior rate that
is (as I understand it) assumed to exist, with satiation effects shown
at high reinforcement rates:

*
* "satiation"
* |
*Behavior rate *| *
* * |
* |
* * |
* |
* * |
* |
* * |
* |
* * reinforcement rate
* *****************|*******************

To avoid satiation effects, one would keep the rewards small enough and
the ratio requirement high enough to avoid getting into the region where
an increase in reinforcement rate no longer produced an increase in
behavior rate.

If we make the reinforcers even larger or the ratio requirement even
smaller, we get into even higher reinforcement rates, so the total curve
now looks like this:

              "satiation"
* |
*Behavior rate |
* *| *
* * | *
* |
* * | *
* Operant cond. |
* * | control *
* region | region
* * | *
* |
* * reinforcement rate *
* *****************|**************************************************

To the right of the satiation point we have "supersatiation." Greater
reinforcement goes with less behavior. The data to the left of the
satiation point follow the expected relationships of operant
conditioning: greater reinforcement goes with more behavior. We have
reason to believe that left of the satiation point, the observed
relationship is a matter of dividing time between the particular
behavior being measured and other behaviors. To the right, we are seeing
essentially continuous engagement in a single behavior and control of a
single variable, misnamed the reinforcer.

I think you're mixing apples and oranges here. A plot showing satiation
effects would give the rate of responding as some function of the quantity
consumed, NOT the rate of consumption as shown in the diagram. Responding
at a rate to the right of your "satiation" line does not mean that the
animal is satiated; if that were true there would be no responding at all in
this region.

In PCT terms we could equate satiation with the reduction in error in the
nutrient system (although again things get complicated when you begin to
consider the different control systems which may be involved, e.g., stomach
loading, blood glucose level, etc.). This lowering of error would reduce
the reference value for food consumption and thus the rate of lever-pressing
on which the food is contingent. I don't think you wish to assert that the
reference level for food consumption (and thus for lever pressing) is zero
to the right of your "satiation" line. Consequently, I'm mystified as to
what you do mean by this line.

The data sets you reported, at least in terms of average rates, fit the
curve to the right of the satiation point. Clearly, the apparent
satiation point found when approached from the left is very much less
than the reference level of the control system, which is all the way to
the right. If you have been identifying the satiation point with the
reference level, perhaps you need to reconsider.

What an odd thing to say. I believe it was I who pointed out in my initial
graph of these data that the reference point, as determined by the straight
line fit, is around 460. I think I made it absolutely clear that I identify
the reference level exactly where you identify it, not at your so-called
"satiation" point.

In reinforcement-theory terms, the region to the left of this point is where
the long delay to reinforcement and the high response cost weaken the net
reinforcement for responding to the point where the "reinforcer" can no
longer support the response, especially if there are alternative, competing
sources of reinforcement available to the animal. I'm working (slowly)
toward a model that will (I believe) demonstrate these effects.

Bill Powers (950630.1130 MDT) --
    Bruce Abbott (950627.1055 EST)

Even if the subtractive cost is a quadratic function of the behavior
rate, the curve never falls below horizontal. It's just the way the
feedback effects work. You need to give the cost-benefit variable a
nonlinear effect on the control system itself such as varying the gain
in its output function. Then you can get the two-valued effect. I've
tried most of these variations, and there are probably still others I
haven't thought of.

Then by adjusting the gain of this loop you can get the
whole curve, over ratio requirements from 1 to 160, to match the real
data very closely.

    It seems to me that this solution is ad hoc. Why would one expect
    the reference level for effort to be set so high?

Yes, quite ad hoc. It's the only ad hoc model I found that worked. One
reason why the reference level for cost might be set high is that below
some amount of effort, the body's normal resupply control systems can
keep up with the energy expenditure, so in effect the only net cost is
in use of stored energy which would be wasted anyway if not used. As the
food supply diminishes and the efforts required to maintain it increase,
there comes a point where normal metabolism starts to fall behind the
rate of energy usage, and that is where I would expect the commencement
of attempts to conserve energy by reducing activities -- lowering the
gain in many control loops. High-gain control loops simply consume more
energy than low-gain loops, because of correcting tiny errors all the
time.

In other words, local cost can exceed local benefit by a certain amount
as long as the whole system can make up for the losses. When the whole
system reaches its control limits, we would expect some sort of major
adjustment to begin.

O.K., that sounds reasonable, if not convincing. My guess (also only a
guess) is that it doesn't take such a crisis to get a rat to abandon
unproductive (or counterproductive) activities. At some point the gain is
just not worth the cost. But good, perhaps we can find a way to test these
alternative hypotheses.

Regards,

Bruce

[From Bruce Abbott (950701.1355 EST)]

Bill Powers (950630.1930 MDT) --

OK, I see that I had a false impression from the start. When, in the
first post, you said "here's some data we didn't have ... before showing
a second set of figures, I took it to mean that this was a different
experiment. Now I understand. Sorry if I'm being dense.

Sorry if that confused you. My comment about data we didn't have was in
reference to data not presented in the Motheral graphs. Motheral's graphs
showed overall rates of responding and reinforcement but did not provide a
breakdown of theose overall rates into the component postreinforcement
pauses and running rates.

By "satiation" I meant to indicate the boundary between the region where
increased reinforcement would produce increased behavior, and the region
where it would not. I presume that in an experiment where it is expected
that increasing the reinforcement would increase behavior, and there is
a falling off of the rate of increase or a leveling out of the amount of
behavior with increased reinforcement, the experimenter might conclude
that satiation is being approached, and either reduce the amount of
reinforcer or increase the ratio requirement to keep this from
happening. If this had been customary, it might explain why the region
to the right of the line was not explored by the early experimenters.
And that, in turn, would explain why the effects of reinforcement are
traditionally described as they are.

But this is not at all the meaning of the term as used in conditioning
experiments. "Satiation" means just what it means in ordinary discourse.

Operant conditioning folks would consider the right limb of Motheral's curve
the region of normal operation for ratio schedules; what happens on the left
side is called "ratio strain" and is said to represent a breakdown of
"control" over the operant by the schedule. In this region, as the ratio
requirement increases, the reinforcement is simply too weak and instability
sets in: the animal may quit responding altogether, or may return to the
lever sporadically.

The basic difficulty in understanding this situation, I think, is that one
must distinguish two roles of the reinforcer. The first role is as a
promoter of learning, which in this context is basically the selection of
behavior: what to do. In this role, reinforcement is said to "strengthen"
behavior which leads to reinforcing consequences.

The second role is as the source of motivation for executing the behavior
which has been learned. The behavior may have been "selected" (first role),
but if there is no motivation it will not be executed. In the steady state
(when "selection" has been completed), the rate of behavior will vary with
the level of motivation.

During training, behavior rate may be low because (a) the response has not
been completely "selected" and/or (b) the motivation to respond is low.
During steady-state performance, "selection" is complete, so the only
influence on performance is the level of motivation.

At present, PCT and reinforcement theory make contact primarily in the
latter domain. And it is here where you keep attempting to apply
reinforcement theory's learning function to explain steady-state responding,
without considering the motivational role which reinforcement and punishment
are assumed to play in the steady state. For this reason I view some of
your attempts to apply reinforcement theory to steady-state conditions as
completely off the mark.

So let's leave out selection for now in order to focus on the motivational
aspects of reinforcement, and see how they apply. Assume that a rat has
been trained to press a lever to earn food pellets on a variable-ratio
schedule (say, VR-20). If there were no other sources of reinforcement in
the environment, and if lever-pressing had no cost, one might expect the rat
to press the lever at maximum rate so as to maximize the rate of
reinforcement, regardless of the level of motivation. After all, there's
nothing else worth doing, no cost to responding, and however tiny the
reward, it beats no reward. But these are ideal conditions never found in
nature or even in the laboratory.

In reality every behavior has a cost function associated with it. First,
there is the pure expendature of energy, and the wear and tear on the
muscles and joints. Second, there is the time such efforts take away from
doing other things. That the rat does other things between reinforcements
besides pressing the lever indicates that there are other sources of
reinforcement which serve to motivate those behaviors (otherwise they would
not occur, right?). To some extent these other behaviors and lever-pressing
can be "time-shared" or interlaced; however, as the rate of lever-pressing
increases, less and less time is left within which these other behaviors may
occur. Similarly, the stronger these other sources of reinforcement, the
more time the rat will spend engaging in other behaviors and the less time
will remain for lever-pressing.

The rate of lever-pressing observed on the VR schedule will thus represent a
dynamic equilibrium involving the reinforcing value of the reward, the
response cost, the value of competing sources of reinforcement, and the
degree to which lever-pressing and the other behaviors motivated by those
competing sources of reinforcement interfere with each other.

Now, what determines the reinforcing value of the food pellets, that is, the
effectiveness of the reinforcer in maintaining the behavior that produces
it? Again the situation is somewhat complex. All else being equal, the rat
will respond more vigorously for a pellet of food the higher the level of
deprivation (up to the point where deprivation begins to weaken the animal,
a region rarely explored in operant studies). So higher deprivation leads
to the pellet being more highly valued, capable of sustaining a higher
response rate. The size of the pellet has a similar effect (note: both
effects are nonlinear. Each successive doubling of pellet size does not
double the value of the pellet; a log function may be indicated.) The
flavor of the pellet would also be a factor determining value, and several
other things as well that I will skip over in the interest of brevity.

In addition, there is contiguity between the response and the delivery of
the pellet to consider when determining value. The larger the delay between
responding and reinforcement, the weaker the effect of the reinforcer on the
response (the so-called gradient of reinforcement). Everything else being
equal, a response which is followed by a reinforcer 5 seconds later will not
be sustained at as high a rate as one which is followed immediately by the
same reinforcer.

All these relationships I have been describing have been empirically
measured; they are not simply invented as needed, with just the right
properties, to make reinforcement theory "work" in a given situation. To be
fair, you must distinguish between ad hoc hypothesis which have been offered
at the end of a study to explain a result (offered as hypotheses to be
tested in future research, not as the final word) and explanations based on
factors established through empirical testing.

When one asserts that reinforcement theory "cannot" explain some finding,
you are inviting the reinforcement theorist to generate an ad hoc
explanation which demonstrates that there are ways in which the theory _can_
explain the finding. But such an explanation is not the kind a competent
reinforcement theorist would offer if asked to show that reinforcement
theory _does_ explain the data. The latter type of explanation would have
to include the results of empirical tests which support that explanation.

There are other factors which have come to light over years of empirical
research that also need to be taken into account when developing a
reinforcement model of behavior in a given experimental setting (e.g.,
effects that depend on the animal's ability to discriminate temporal
intervals, so-called "conditioned" reinforcement, presence of
discriminative stimuli, and so on), so the application of reinforcement
theory as currently developed is not simple when several of these factors
appear to exert an important influence. The same will be found when PCT
models need to be developed to handle these same empirical relationships:
those models will have to include a number of controlled perceptions whose
control systems interact. However, these models will be simpler than the
reinforcement models, because effects that emerge naturally from the
structure of the control system must be referred to some complex
interactions among factors of the reinforcement model (i.e., computations of
eccentrics and epicycles).

The total quantity consumed, it seems to me, would be meaningful only if
the time over which it is consumed is also stated. After all, we have
all eaten thousands of meals without becoming permanently satiated.
There is a balance between intake and metabolism such that the internal
nutrient levels may be higher or lower depending on the rate of intake.

It's funny: I actually erased a paragraph saying this in my post. I deleted
it because it seemed to distract attention from the main point I was trying
to get across. The rate at which satiation effects would appear would
depend on the rate of food intake and on the rate at which the food eaten
was consumed by the rat's bodily processes. But if the vertical line on
your graph labeled "satiation" represents the point at which intake exceeds
consumption, how does this relate to the rates observed to the right of the
line? Satiation is not a rate but a state, and if you have reached
satiation, there is no motivation for behavior whose consequence is to
produce food.

Perhaps I have simply misunderstood how satiation enters into
explanations of failures of reinforcers to reinforce. If this is _not_
the explanation that would be used in the case before us, how _would_
the behavioral theorist explain the decrease of behavior rate that goes
with the increase in reinforcement rate to the right of the line? My
impression is that behaviorists have described the relationship between
behavior and reinforcement that we can see to the left of the line; I
vaguely recall some admonitions by my professors that one should not
make rewards so large that behavior reaches an upper limit and levels
off. They never mentioned the possibility that it could actually decline
with further increases in reinforcement. I think I remember satiation
being mentioned as a reason for the leveling off.

The problem your professors admonished you about is that if rewards are
large enough you will reach the upper end of the rat's response rate. If
the rat is already responding at a maximal rate, your dependent variable
(response rate) has nowhere to go but down--a ceiling effect. If you are
going to manipulate other parameters, you would rather have a dependent
variable that is free to change in either direction. A reinforcer can't
"reinforce" (increase rate of responding) if the rat's rate of responding is
already at maximum.

This problem has nothing to do with satiation.

Regards,

Bruce