Ratio Model

[From Bruce Abbott (950626.2200 EST)]

Between starting summer school teaching and getting out book manuscript I've
been a little bogged down today, but here's my two cents...

Bill Powers (950623.1410 MDT) --

OK, so you have to make the behavior rate depend on the value of the
reinforcer to the subject, minus the cost of responding. How do you
introduce the value of the reinforcer to the subject in the model? Does
this require adding something to the model equivalent to a reference
signal?

Yes, but a better equivalent would be measuring the relative size of the
error signal. The "value" of a reinforcer basically is a measure of its
motivating properties under given conditions. You can measure this value
by, for example, determining the rate of responding that a given reinforcer
will sustain under given conditions, e.g., a given quantity of a specified
food under 12 hrs of food deprivation. A larger quantity, better taste, or
greater hours of deprivation would be expected to yield higher rates,
indicating that the reinforcer is more highly valued under those conditions.
These operations increase the size of the error signal either by changing
the reference level or by applying a larger disturbance.

    Let's apply this model to the steady-state situation involved in
    performance on a simple CRF (1 response per reinforcement) schedule
    as ordinarly studied in the operant chamber. A hungry rat is
    placed in the chamber. We might conceive of the set point for rate
    of food pellet consumption as, say, 30 pellets per minute (30 ppm)
    under these conditions (essentially continuous eating). However,
    the apparatus limits the maximum rate to 10 ppm because of the
    delays involved in depressing the lever, moving to the food cup,
    picking up the food, devouring it, and returning to the lever.
    Thus, there is NO WAY that the rat can reach its reference level
    for this quantity, although it can reduce the error by a
    significant amount through lever-pressing (and by no other means).
    I believe that control theory predicts that the rat (once it has
    learned what to do) will develop a rate of responding that
    minimizes the error, up to the point where the error is reduced
    enough to bring the output below maximum. Depending on the gain,
    there will be a region within which the rate of responding will be
    a direct function of the magnitude of the error.

All right, you're setting up a special situation in which the rat can't
reach its reference level for nutritive input because there are physical
limitations, combined with the reward size, that prevent its doing so.
The exact prediction that a control model would make would depend on the
loop gain of the rat (measured without these restrictions), on the
measured reference level for food intake, and on the degree of the
restrictions. If the rat had a high loop gain, we would expect the
behavior to remain at the maximum rate under all conditions.

Let's graph the situation:

      450*
         *
         * x
         *
         * x
      300* - - - - - +
BEHAVIOR * | x
         * |
         * | x
RATE * | ref level
      150* | x <--error--> |
         * | |
         * | x |
         * | |
         * | x |
         *************|*************|************|**********
                      10 20 30
                   REINFORCEMENT RATE

The line of x's represents the minimum consistent output sensitivity of
the control system in units of presses per minute per reinforcement per
minute of error. To fit your conditions it has to be about 30
(reasonable).

    If we now increase the ratio requirement, the maximum rate of
    reinforcement available on the schedule is less (say, 5 ppm).

It would have been easier if you had specified the new ratio requirement
and then figured out the result, but we can still get there. If the new
rate of reinforcement is now 5 reinforcements per minute, we can deduce
that the new ratio is 390/5 = 78.

      450*
         *
      390* - - x
         * |
         * | x
      300* |
BEHAVIOR * | x
         * |
         * | x
RATE * | ref level
      150* | x <--error--> |
         * | |
         * | x |
         * | |
         * | x |
         ******|********************|************|**********
               5 10 20 30
                   REINFORCEMENT RATE

I've been puzzling over these two graphs for several days now, trying to
figure out where I'm getting lost. I think I now see the problem. Your
graph has little to do with my scenario, and it's got an error in it to boot.

First, the error. If you are using an output sensitivity of 30, then in the
first graph, the error between 30 rft/min and 10 rft/min is 30-10 = 20
rft/min, which would be expected to generate 30 * 20 = 600 responses/min and
not the 300 you indicate in the figure. Right?

Second, the scenario. In my off-the-top-of-my-head example, I suggested
that the maximum response rate on the CRF schedule is 10 rft/min, including
the time required to collect the food, consume it, and return to the
response lever. For the sake of illustration, let's assume that pressing
and releasing the lever consumes 0.2 sec per press. Now, 10 rft/min is
equivalent to one "cycle" of press-consume-return per 6 sec. If 0.2 sec is
required for the press, this leaves 6.0 - 0.2 = 5.8 sec for the remaining
activity. And at FR-1, 10 rft/minute = 1 * 10 = 10 ppm (presses per minute).

At the unspecified higher ratio, I stated that the maximum rate would be 5
rft/min, or 12 sec per cycle. With the consummatory activity still
requiring 5.8 sec to complete, this leaves 12.0 - 5.8 = 6.2 sec for
lever-pressing, or, at 0.2 sec per press, 31 presses. Thus (under the
assumption of 0.2 sec/press) the new ratio is VR-31. At VR-31, 5 rft/min
translates into 31 * 5 = 155 ppm (presses per minute). Remember, this is
the _maximum_ rate at which food can be collected and eaten. So, in my
hypothetical set of observations, the rat cannot reach its reference rate of
food delivery on the CRF schedule (error = 20 rft/min) and does even worse
under the new schedule (error = 25 rft/min), when responding at maximum
rate. Yet there is an inverse relationship between rate of reinforcement
and rate of responding, simply because of the extra time required to
complete the ratio as the ratio requirement is increased.

Let's graph the situation:

···

*
          *
       150* +<----------------- error ----------//--->|
          *
          *
          *
          *
       100*
BEHAVIOR *
          * +<---------------- error ----//--->|
          *
RATE *
        50*
          *
          *
          *
          * +<--------- error ---//--->|
          *************|*************|************|*******//****|
                       5 10 15 30
                      REINFORCEMENT RATE

The error is increasing with the ratio requirement, but behavior is already
at max rate and so can't get any faster. The line defined by the + symbols
represents the constraint imposed by the time required to earn and collect
the reinforcer, and not any regulatory effect.

Now, assume that the reference is 10 rft/min rather than 30, and that the
output sensitivity is 30 as you suggest. On CRF the rat can earn its
reference food rate by responding at maximum ppm. Assuming no constraints,
at 5 rft/min the error is 10 - 5 = 5 rft/min, which translates into 5 * 30 =
150 ppm, giving a schedule value of 150/5 = VR 30. The maximum rate on this
schedule is 152.5 ppm. Up to this point, at least, the schedule constraint
does not enter in and the negative slope of the function is due to regulation.

I have more but it'll have to wait ... time's up.

Regards,

Bruce

[From Bruce Abbott (950628.0900 EST)]

Bill Powers (950628.0400 MDT) --
    Bruce Abbott (950627.1645 EST)

More good data, there. Did the authors make any comment at all about the
fact that more reinforcement goes with less behavior?

    The inverse relationship between the overall response rate and the
    schedule parameter has also been observed on random-ratio schedules
    (Brandauer, 1958, Kelly, 1974). These functions would appear not to
    be consistent with the Law of Effect (Herrnstein, 1961; 1970), as
    the Law would predict an direct and not an inverse relationship
    between response rate and rate of reinforcement, as determined by
    the schedule parameter.

    (Priddle-Higson, Lowe, & Harzem, 1976, p. 353)

However, these authors felt that these results might be explained by the
complex interaction between running rate and postreinforcement pause
observed in the data. The overall rate would result from the combination of
these two factors.

Average 3.9 2.7 2.7 rsp/sec 08 10 15 seconds
          234 162 162 rsp/hr

This is clearly rather extreme behavior: 14000 responses per hour! I don't
know what your last line above means -- it's obviously not 3600 times the
line above it.

Sorry, I had been working with rsp/min and later converted to rsp/hr for
comparison with Motheral's data. I forgot to do the conversion on the
running rate data, as I did not plot them. So the numbers in the lower line
are rsp/min. 14000 rsp/hr is a high rate, but that's not the sustained
output over the session. On ratio schedules, once the rat begins the ratio,
it nearly always knocks it off in one continuous series at a steady (and
fairly high) rate.

The running rate reflects the time required to complete the ratio once the
rat has resumed responding following the postreinforcement pause and is not
the rate maintained overall. In my hypothetical scenario I had suggested
0.2 sec per response as a reasonable figure for the maximum rate possible,
which gives 5 responses per second, compared with the 3.9 shown in these
data. The highest running rate shown in the data presented in the article
is around 5.2 rsp/sec. This rate was recorded for one animal on the VR-10
schedule at a 10% milk concentration, the lowest tested.

I would guess that this set of runs represents a very dilute solution of
condensed milk; the curves seem to be on the border between positive and
negative slopes. My hypothesis is that you get a positive relation between
reinforcement and behavior when not enough reinforcement is recieved to
support life if continued indefinitely.

The condensed milk was diluted to a 30% concentration, or about 2 parts
water and 1 part concentrate. This is not especially dilute. The dipper
presented 0.05 ml per dip, a medium-sized raindrop.

That's my hypothesis, too. However, you do have to keep in mind the fact
that these rats will be given supplemental food after each experimental run,
so a large deficit is not going to appear over the long run as it would if
all the rat's nutrient intake had to come from lever-pressing. Given the
actual arrangement, it is possible to induce a rat to work at a (short-term)
loss by increasing the ratio requirement slowly enough, although even then
there are limits beyond which the behavior will break down.

There is one problem with using dilution as a way of altering the amount of
reinforcer. The less the food content, the more water the animal has to
drink to get the same amount of food reinforcer. So the water-loading
control system is going to have a "too-much" error, and start fighting the
food control system.

Yes, that could be a problem; on the other hand the stomach loading is
constant on a per-reinforcer basis when the amount is varied in this way, so
there are advantages as well as disadvantages to this approach. The
water-loading control system has a simple way to compensate for the error
other than reducing the intake rate: excrete the water faster. So perhaps
this is not as serious a problem as it may seem to be at first blush.

I think that when we analyze data we should do it on a rat-by-rat basis,
not average the data across rats. In the above data set, the rats don't all
behave alike. Better to apply the model to each rat, and then compare the
parameters.

I agree, and that is why I presented the individual data in my post. One of
the difficulties in gathering data like these is that it is not always
possible to keep variables such as motivational level (error) perfectly
constant over the course of an experiment. For this reason some of the
individual data may not be as consistent as one would like. In the data
presented, most subjects showed similar changes in rates across the
schedules, but there were exceptions. The averages I graphed seemed to be
representative of what most subjects did most of the time.

It's interesting that using variable ratio schedules seems to give the same
kind of results that Motherall got with fixed ratios.

Yes, although I'm not completely certain that Motheral used fixed ratios
(Staddon keeps referring to "ratio schedules" without designating whether
they were fixed or variable ratios). But if so, the similarity of results
with fixed and variable ratios would have emerged despite differences
between the two schedules in the local patterns of responding associated
with them.

I haven't taken a look at your program yet, so I won't comment on it now.

Rick Marken (950627.2100) --

Very interesting data, Bruce. But could you explain "running rate" in
a little more detail. It sounds like "running rate" is the response rate
that occurs after the animal has worked through the ratio requirement.
Is that it?

It's the response rate observed WHILE the animal is working through the
ratio requirement, timed from the first response following the
postreinforcement pause.

Also, what were these researchers trying to find out by collecting these
data? I presume they were testing some predictions of a reinforcement
model? Did they compare their data to the data produced by a working
reinforcement model? Were these data consistent with the predictions
their model?

In the experiment I described the researchers were assessing the effect of
milk concentration on the length of the postreinforcement pause, on the
running rate, and on the overall rate of responding, in the interaction of
these effects with FR ratio size. The study does not appear to have been
designed to test predictions of a reinforcement (or any other) model,
although there was some discussion of the anomaly apparent in the data with
respect to reinforcement theory predictions. Rather, it was designed to
provide information about the relationships being examined (Baconian style).
They did not develop a working model.

Regards,

Bruce

[From Bruce Abbott (950628.2020 EST)]

Bill Powers (950628.0945 MDT) --
    Bruce Abbott (950628.0900 EST)

    However, these authors felt that these results might be explained
    by the complex interaction between running rate and
    postreinforcement pause observed in the data. The overall rate
    would result from the combination of these two factors.

This is a pretty feeble comment on an observation that goes directly
against the fundamental assumptions behind reinforcement theory itself.

Well, these guys aren't reinforcement theorists, they're just humble
experimentalists reporting their data. I give them credit for perceiving
and pointing out the apparent problem for reinforcement theory; apparently
they felt less than comfortable going much further. Their objective was to
discover what relationships occur in this situation, and that's what they
did. They were willing to leave the theorizing to someone else.

When you compute the _average_ values of reinforcement rate and behavior
rate, the same relationship is seen. In fact, the equivocal slopes
become the "wrong" slopes for reinforcement theory. (You can get the
average behavior rate by multiplying the average reinforcement rate by
the ratio).

Huh? The _same_ relationship when you compute the average values? The
average, overall rates are precisely the ones the researchers were talking
about when noting the inverse relationship, so what other relationship can
you be speaking of? I'm confused.

What the researchers were noting is that the overall response rate can be
viewed as being determined by two factors: the running rate and the
postreinforcement pause, and that these changed in different ways with the
ratio requirement. In some cases these changes tended to oppose one another
in their effect on the overall rate; in others they tended to summate.
Given that the overall changes reflect an average of these more local
changes, then one might suspect that the overall rates obfuscate the actual
processes at work. That seems reason enough to be at least a little careful
about announcing to the world that your data absolutely contradict a
reinforcement analysis. If you found some rather odd things happening in a
complex tracking task that seemed to contradict PCT, you might note that the
results don't appear to agree with the PCT prediction but I doubt you'd be
ready to chuck the whole theory.

The attitude displayed by the authors toward this finding is interesting
from the standpoint of a study in steadfast faith. "No, madame, there is
no danger. A little bit of floating ice could never damage the Titanic."

No, it's more like "there's something odd going on here, but the details are
rather complex so I'm not exactly sure how to interpret these results. I'll
leave that to someone else."

    In the experiment I described the researchers were assessing the
    effect of milk concentration on the length of the postreinforcement
    pause, on the running rate, and on the overall rate of responding,
    in the interaction of these effects with FR ratio size.

That is what they _thought_ they were assessing. Actually, they were
simply looking at the _relationship_ between milk concentration and
those other variables. To say they were assessing the _effect_ of milk
concentration is to assume that milk concentration is _a priori_ a
causal variable.

Oh, come ON! If I can't say that a disturbance to a control system under
given conditions _causes_ it to respond with an opposing action, then I
don't know of _any_ circumstances in which the term would be appropriate.
There is directionality to this relationship; certainly the changes in
response rate, reinforcement rate, and postreinforcement pause did not cause
the milk concentration to change. The word "cause" implies this
directionality in a way that "relationship" does not. You may be able to
rearrange F = MA or I = E/R and get sensible results, but this is different.

RE: water loading

    The water-loading control system has a simple way to compensate for
    the error other than reducing the intake rate: excrete the water
    faster.

I think we have to distinguish "appetitive" control systems, which
control the immediate effects of ingested food or water, from "chronic"
control systems, which work on a much slower time scale and control
longer-term average effects. If you load the stomach with 0.05 ml of
water every 8 seconds or so, that is almost half a milliliter per minute
or about 20 ml per hour -- a lot for a rat's stomach. If this "simple
way to compensate" is seriously offered as reason to ignore the
confounding of food intake with water intake, then the experiment should
be done to verify that this is justified.

I agree. But I'm sure these investigators kept a record of the response
rates throughout the session and would have reduced the size of the droplet
or the session length if there were evidence of satiation. Perhaps that's
why they started with VR-10 rather than something lower. As to time-scale,
I would bet that a rat's kidneys can do quite a bit of error reduction
within the span of an hour.

    One of the difficulties in gathering data like these is that it is
    not always possible to keep variables such as motivational level
    (error) perfectly constant over the course of an experiment. For
    this reason some of the individual data may not be as consistent as
    one would like.

A more likely explanation in my view is that there are individual
differences between rats. While an individual rat might behave in a way
perfectly consistent with a model, the appropriate parameters might be
quite different from one rat to another.

If parameter adjustments can allow a single model to account for the
conflicting changes seen across ratios in some of these animals, then I'm
worried about the model--it would appear to be able to account for any
pattern whatever. When you have that, you're doing curve fitting.

I don't think we would want to keep a rat's "motivation level (error)
constant" in an experiment, even if we could. If you want to measure the
characteristics of a control system, you have to let it control, and
this means letting it correct its errors. The only way to keep error
constant would be to use an external control loop that is stronger than
the rat's control system; this would not reveal normal control
characteristics and might well cause reorganization to start, making all
measures invalid.

Depends on which level you're talking about. Generating a given level of
error in the nutrient-level control system simply establishes a given
reference level for lower-level systems such as the one governing rate of
eating. For the purpose of studying the relationship between lever-press
rate and eating rate as the ratio requirement is varied this may be
perfectly acceptable procedure. This is the system I'm measuring the
characteristics of, and I'm letting it control.

The greatest percent differences were in the data set where the peak
rates of responding were near 14000 per hour, and the average rates
around 6000. That is in the upper left corner of the data plots where
the slope of the relationship may be on the point of reversing. This is
not the best place to evaluate the normal relationships among variables.

Not in the data I'm looking at. The peak (running) rates near 14000 per
hour occurred on the VR-10 schedule, which is at the lower right end of the
curve. The peak (average) rate (5760) occurred on the VR-80 schedule, which
is the leftmost point, but this rate is only slightly higher than the rate
observed on VI-40 (5400).

    The study does not appear to have been designed to test predictions
    of a reinforcement (or any other) model, although there was some
    discussion of the anomaly apparent in the data with respect to
    reinforcement theory predictions.

Whether it was designed to do this or not, it was a test of the
reinforcement model. Every experiment is a test of the basic model,
isn't it?

Yes, but not every experiment yields an unambiguous test of theory, and such
ambiguity is especially likely when the experiment was not designed with
theory testing in mind. Some studies are just carried out to discover
empirical relationships (as I think this one was). After the data have been
collected and published, the theorists can then step in and squabble over
the implications. (;->

Regards,

Bruce

[From Bruce Abbott (950629.1530 EST)]

Bill Powers (950629.0715 MDT) --

Let's massage the data from the second set a
bit more and see what the actual numbers are:

. . . [computations based on observed prp and running rate]

The "only slightly higher" looks quite similar to the same "only
slightly higher" in your original graph of the first data set, which
used the running response rate rather than the mean response rate. It
would be interesting to see the same treatment of the data for the first
set.

I think you're missing something here. The "data for the first set" did not
use running rate rather than the mean response rate, they used mean response
rate. What you've done here is to compute the mean response rates based on
the mean running rates and postreinforcement pause lengths. Here are the
results from the two sources compared (averages only, for illustration):

                Response Rate Reinforcement Rate
          VR-10 VR-40 VR-80 VR-10 VR-40 VR-80
Overall 2880 5400 5760 288 135 72
Formula 3114 5868 6462 311 146 81

"Overall" is the overall rate estimated from the graph and "Formula" is the
overall rate computed from the formula and based on the numbers for running
rate and prp estimated from the graph. The two sets of numbers should
agree; the difference is due to my having roughly estimated the numbers from
the graph presented in the article.

average response rate: responses per unit of time across session
         running rate: responses per unit of time from 1st response after
                            reinforcement to completion of ratio
postreinforcement pause: time from completion of ratio to 1st response after
                            reinforcement

The problem is there are are two independent ways in which a pause can
be created. One is the post-reinforcement pause just noted. But the
other is the possibility of the animal leaving the control lever and
searching for food elsewhere. In the latter case, assuming no food is
found elsewhere, the main result would be a very large error by the time
the animal got back to the lever, and a high rate of behavior
immediately after behavior starts. This could cause a change in the
ratio of average to peak rates, because of large nonlinearities in the
decay of the perceptual signal.

So we really have two cases: one where the animal is staying at the
lever except for collecting the reinforcer, showing a post-reinforcement
pause due to eating and to any overshoot of the perceptual signal, and
the other due to leaving the lever to search elsewhere for food. With
the instrumentation you are putting into your apparatus, we should be
able to see the difference.

Yes, we'll be able to see that. Cumulative records of responding indicate
that animals remain at the lever, once they return to it, and respond
steadily until they complete the ratio requirement (perhaps this breaks down
on the left side of the curve; we'll have to see, although I'll also take a
look at the data on "ratio strain" to see what they may show. So at least
for the right limb of the curve any searching for food elsewhere must take
place during the postreinforcement pause--and would tend to extend it. I
would expect such alternative activities to increase as the ratio increased
beyond some value where the costs of continuing to respond on the lever
begin to seriously offset the benefits of access to the food. The data do
in fact show that pause length increases with ratio size, consistent with
this hypothesis.

Changing the milk concentration alters the effect of behavior rate on
the rate of consumption of the food component of the milk, and also of
the water component. If there is no behavior, however, the milk
concentration has no effect at all on the organism. In fact, the rate of
food consumption is the rate of behavior times the concentration, so it
is possible that changing the milk concentration will have no effect at
all on the rate of food consumption, or that the rate of consumption can
rise or fall with no change in the concentration.

Yes, and holding your hand at right angles to the airstream will have no
effect at all on drag if you aren't moving against a relative wind, either.
Yet under specified conditions: e.g., experiencing a 50 mph relative wind in
my face, I think it is quite appropriate to say that the position of my hand
influences the drag. Where behavior influences perception which influences
behavior this notion of directional influence loses meaning, but I see
nothing wrong with such statements as, "given a particular reference level,
adding disturbance X to the controlled perceptual variable causes the system
output Y to change." Here X influences (via the control system's structure)
Y, but Y does not influence X; the influence is clearly directional. The
linkage is via the control system, so changes in the system will alter the
linkage, as when the reference changes.

    But I'm sure these investigators kept a record of the response
    rates throughout the session and would have reduced the size of the
    droplet or the session length if there were evidence of satiation.

Oops. Two problems. First, the _entire_ data set number 1, and most of
number 2, shows effects of satiation -- it's over the top of the curve.
And second, this is precisely what I was talking about a while back when
I said that the standard conditions used in EAB experiments were
designed to avoid satiation (coming up the Motherall curve from the
left) and thus never reached the conditions where the control phenomenon
would be seen. That's where the standard concept of reinforcement comes
from. In the present experiments, the standard concept of reinforcement
doesn't even apply.

You'll have to explain this to me. On what basis do you say the "entire
data set ... shows effects of satiation"? Second, if we're investigating a
hierarchical system, how does keeping the reference level for one system
raised to a particular value (by means of keeping a given error present in a
higher-level system) prevent control phenomena (in the lower-level system)
from being seen?

Regards,

Bruce

[From Bruce Abbott (950629.1730 EST)]

Rick Marken (950629.0910) --

Bruce:

Oh, come ON! If I can't say that a disturbance to a control system
under given conditions _causes_ it to respond with an opposing
action, then I don't know of _any_ circumstances in which the term
would be appropriate.

Oh, come ON, Bruce! The term "disturbance" is not a PCT replacement for the
word "cause". Disturbance refers to a variable that causes another variable
to be moved from (or to) a _preferred state_; a disturbance variable is only
a disturbance to a _controlled variable_; otherwise, a disturbance is just a
variable that causes a change in another variable.

Now, Rick, I didn't say that "disturbance" is a PCT replacement for the word
"cause," I said that a disturbance IS a cause. You just said this yourself:

Disturbance refers to a variable that causes . . .
a disturbance is just a
variable that causes a change in another variable.

I rest my case.

You said "the researchers were assessing the effect of milk concentration on
the length of the postreinforcement pause". This clearly implies that the
researchers thought that milk concentration might have a direct effect, via
the organism, on an aspect of responding (postreinforcement pause).

How else would it have any effect an aspect of responding but through the
organism?

If you
want to play the PCT translation game you have to play fair; if you want to
call milk concentration a disturbance, then you must show that this variable
has an effect on postreinforcement pause via the joint effect of both
vaiables on a controlled variable: you must descbibe the controlled variable
and explain how milk concentration and postreinforcement pause affect it.

Well, if changes in milk concentration were not a disturbance, why did the
behavior change, Mr. PCT theorist? You're not suggesting that milk
concentration directly affects some reference value are you? Or that the
change in behavior along with milk concentration was coincidental? What
else is left?

The results of my little experiment did not threaten the basic tenets of PCT
(as ratio data threaten the basic tenets of reinforcement theory) but they
certainly were not what was expected, so they demanded explanation. That's
how you feel about deviations from prediction when you are used to a model
of living systems that predicts correctly every time.

Let's see... the results were incompatible with the prediction but were not
incompatible with the theory.... Sounds like what those EAB researchers
must have thought. The results were not as predicted, and so demanded an
explanation. That's what those researchers thought... The model predicts
correctly every time. But it doesn't predict correctly every time, because
it predicted the wrong result in this case. Hmmmmm.... Perhaps PCT
[substitute reinforcement theory] is correct but the particular model based
on it was wrong, thus generating incorrect predictions. So the PCT
[reinforcement] theorists needed to find a better model.... Somehow, I'm
just not getting the difference between your behavior and that of EAB
researchers. (;->

So here was a result that seemed to contradict a prediction of PCT. And it
caused considerable concern and interest until Bill Powers realized (and
showed) that the result is predicted by a control model with a transport lag
(all real control systems have some transport lag).

And you assert (without having done a review of the literature in this area)
that EAB researchers did not come up with a model. Hey Rick, when no
evidence is sought, lack of evidence is not evidence.

So where has that someone else been? It's been nearly 20 years since these
results were reported. And these results are not an isolated case; as Bill
noted, Skinner and others have collected tons of "scheduling" data that show
rather conclusively that consequences don't strengthen responses. Shouldn't
the theorists have shown by now that these results are consistent with the
idea of selection by consequences -- if they are? It looks pretty fishy to me.

Perhaps they have, have you looked? I believe your form of argumentation is
called "begging the question."

Regards,

Bruce

[From Bruce Abbott (950703.1025 EST)]

My last post had to be ended before I had a chance to deal with some
methodological issues Bill P. raised. So, to continue...

Bill Powers (950702.0500 MDT) --

The main problems with your narrative descriptions are (1) that they are
not mathematical, and (2) that they are piecewise.

I have been providing narrative descriptions because I believe that it is
necessary to establish a basic understanding of concepts before we get into
the mathematical modeling. I don't think I would have gotten very far
attempting to understand control theory if I had started with a complex PCT
model rather than the nice descriptions you took the trouble to provide of
the basic elements of control theory: perceptual variable, input function,
perceptual signal, comparator, reference signal, error signal, output
function, gain, disturbance. The piecewise nature of the presentation
emerges from the necessity to discuss each concept one at a time.

We are working toward developing a mathematical model. I want to be sure
you understand the constructs and, at least in a roughly quantitative way,
the expected relationships, before we proceed further.

When a system of
interacting variables is described in words, the quantitative
relationships among the variables become ambiguous; it is not clear, for
example, which of two opposing effects will win. And when the system is
described piecewise in terms of immediate effects of one variable on
another, there is no way to take indirect reverse effects through other
paths into account. Even if a positive change in variable A is observed
to cause a positive change in variable B, it is perfectly possible that
because of systemic effects, the net result of attempting to increase A
will be a decrease in both A and B. (Perhaps Wayne Hershberger would
oblige the net with a description of his famous experiment with chicks,
in which they had to move away from a target in order to bring it closer
to them). Verbal descriptions, and especially piecewise verbal
descriptions, are unable to reveal such effects. Only a quantitative
mathematical description can enable us to see the implications of a set
of piecewise observations on the operation of the whole system.

I agree. And I've read Wayne's chick experiment (nicely done!). As I
recall, under certain conditions some chicks learned to control by moving in
the "wrong" direction and some steadfastly persisted in attempting to pursue
the target.

The general approach of psychological experimentation is piecewise. That
is, most variables are held constant, one variable is changed, and the
effect on another variable is observed. Then, in a different experiment,
a different variable is manipulated while the former variable is held
constant along with the others. The overall impression one gets out of a
long series of such piecewise experiments is that the entire range of
phenomena has been systematically covered -- that the set of piecewise
facts, taken together, adds up to a complete description of the behavior
of the system.

This approach works fine during the initial phases of the research when one
is simply attempting to identify the effective variables and their
relationships under specific conditions. It is, in fact, the prototype of
the experimental method: vary one thing (independent variable), hold
everything else constant, observe any effect (dependent variable).

The limitation of this approach is well known (even in psychology): the
relationship between X and Y may change depending on the value of variable
Z. For example, Z may turn out to be a parameter whose value acts as a
multiplier (e.g. gain) between the values of X and Y. Such effects are
called interactions.

It is quite possible for the piecewise description to result in apparent
contradictions. The reason is that each component of the system is
actually being described under different conditions. An experiment that
shows that increasing deprivation results in an increase in behavior is
done under conditions where behavior is not allowed to reduce
deprivation -- where deprivation may actually be increasing. But another
experiment which shows that reducing the ratio requirement results in a
reduction in behavior is done under conditions where the behavior can
produce more reinforcement and thus reduce the degree of deprivation.
The immediate relationship under investigation would be affected by
other paths if all variables were free to change. It is the fact that
the other variables are under external control, and are held constant,
that creates the apparent contradictions, as well as the apparent

Yes. It is not unusual for researchers in different labs to carry out
similar experiments involving the same independent and dependent variables
and obtain contradictory results. If the results are contradictory, it is
clear that there must be other variables, differing in value between the two
studies, whose differing values must explain the contradictory results.
Such contradictory findings often lead to research in which the suspected
variables are manipulated to determine the source of the interaction and
thus resolve the discrepancy.

You seem to be saying that in your view, psychology experiments always
manipulate only one variable at a time. I would agree that parametric
studies are not as common as they should be, but there are certainly plenty
of examples of them.

Perhaps the reason they are not more common is that running them often
requires an enormous expenditure of time and, in some cases, money. An
example is provided by the Priddle-Higson et al. (1976) study. During the
baseline phases of the first experiment, each rat was exposed to each of the
three VR schedules for 30 sessions in order to achieve steady-state
performance levels on each schedule. That's 90 days just to complete the
baseline phase, manipulating one variable (the ratio) over only three
levels. At each of these ratios, a second variable (milk concentration) was
introduced (5 concentrations) within-sessions to determine how the milk
concentration presented after completion of a given ratio influences the
pause in responding immediately following the consumption of the milk. It
was found that the postreinforcement pause was longer following the
consumption of more concentrated milk than following the consumption of less
concentrated milk; the increase was to a first approximation a linear
function of concentration. In addition, the slope of this function was
steeper the higher the ratio requirement of the schedule (i.e., ratio and
concentration interacted).

Because the effect of milk concentration was immediate, there was no need in
this study to introduce a new concentration and then wait for stability to
set in over perhaps another 30 days; if there had been such a need, then the
experiment conceivably could have required 3 * 5 * 30 = 450 days to
complete. I once ran a study in which two variables were parametrically
varied, which required over two years to complete.

Of course, stable performance may be reached in far fewer than 30 sessions,
so studies are not necessarily going to be this long (5-10 sessions is
probably more typical), but even here the time required to complete a study
is usually measured in months, not hours or days. One can also save time by
abandoning single-subject methodology. For example, in the Priddle-Higson
experiment one could use 3 X 5 = 15 groups of 10 rats each and then plot
average performance of the groups as a function of the two parameters.
Whether you would be learning anything about how these parameters affect the
performance of single subjects is open to question, but I've seen plenty of
studies like that.

    The problem your professors admonished you about is that if rewards
    are large enough you will reach the upper end of the rat's response
    rate. If the rat is already responding at a maximal rate, your
    dependent variable (response rate) has nowhere to go but down--a
    ceiling effect. If you are going to manipulate other parameters,
    you would rather have a dependent variable that is free to change
    in either direction. A reinforcer can't "reinforce" (increase rate
    of responding) if the rat's rate of responding is already at
    maximum.

    This problem has nothing to do with satiation.

My professors thought it did. If the reward size is too large, the
animal will produce all that it wants by hitting the bar at some
moderate rate -- nowhere near its maximum possible rate of behaving. It
may not be totally satiated, but it's getting close. My professors
assumed that if you continued to increase the reinforcement, the
behavior rate would simply remain the same -- but of course it wouldn't.

O.K., perhaps I misunderstood. You said:

I vaguely recall some admonitions by my professors that one should not
make rewards so large that behavior reaches an upper limit and levels
off. They never mentioned the possibility that it could actually decline
with further increases in reinforcement. I think I remember satiation
being mentioned as a reason for the leveling off.

When you said that "behavior reaches an upper limit" I was picturing a curve
in which behavior rate increases with the size of the food pellet. The
upper limit of this curve would be its asymptote (the curve would be
increasing but negatively accelerated). This upper limit has nothing to do
with satiation. What you are now describing is a lowering of the upper
limit (reduced asymptote) due to a lowering of deprivation.

Regards,

Bruce

[From Bruce Abbott (950702.1520 EST)]

I think we're making progress.

Bill Powers (950702.0500 MDT) --
    Bruce Abbott (950701.1355 EST)

What we're concerned with, then, is the "motivational" aspect of
reinforcement, which covers the right (descending) side of the
behavior/reinforcement curve.

Why just the right side? The motivational aspect should cover the whole
curve, since we're dealing with steady-state behavior.

    The second role is as the source of motivation for executing the
    behavior which has been learned. The behavior may have been
    "selected" (first role), but if there is no motivation it will not
    be executed. In the steady state (when "selection" has been
    completed), the rate of behavior will vary with the level of
    motivation.

This is unclear. If the reinforcement is the source of motivation, then
how does motivation vary with reinforcement rate? Does motivation
increase as a function of increasing reinforcement rate?

If we consider food pellets as reinforcers, then for a given level of
deprivation (held constant!) and a given type and size of pellet, etc., the
motivation attributed to the reinforcer would be constant across rates.
However, responding brings costs which may vary as a function of response
rate. These costs are effectively a source of motivation for not
responding. The observed rate of responding would reflect a sum of these
opposing motivational levels.

Or is
motivation an independent variable determined in some other way?

See below.

Apparently, a reinforcer can be present in the role of "selecting" a
given behavior, without at the same time automatically "motivating" it.

No, to provide selection, the reinforcer must also motivate. However, once
the steady state is reached, there can be motivation without further selection.

    During steady-state performance, "selection" is complete, so the
    only influence on performance is the level of motivation.

What do you mean by "performance?" Behavior rate?

Performance is behavioral output, which would cover all facets, including
such aspects as rate, duration, response force, even temporal pattern.

If reinforcement is
the source of motivation, then can't we just say that behavior rate goes
with reinforcement rate?

No. See above.

Does the term "motivation" have a meaning other
than rate of reinforcement? Is it possible to have high motivation and
low reinforcement rate, or high reinforcement rate and low motivation?
Or is motivation merely, to a first approximation, proportional to
reinforcement rate?

One could have high motivation (say, because of high deprivation) and low
reinforcement rate (because of schedule parameters). If the reinforcer can
be earned with low effort and there are no alternative sources of
reinforcement, fairly high rates might be sustained at relatively low levels
of motivation.

    Satiation is not a rate but a state, and if you have reached
    satiation, there is no motivation for behavior whose consequence is
    to produce food.

Satiation, then, works to reduce motivation: i.e., the closer the
organism approaches to satiation, the less motivation there is. So
motivation depends on a state of the organism that is affected by the
reinforcer but that can vary independently of the reinforcer.

Yes.

    The rate of lever-pressing observed on the VR schedule will thus
    represent a dynamic equilibrium involving the reinforcing value of
    the reward, the response cost, the value of competing sources of
    reinforcement, and the degree to which lever-pressing and the other
    behaviors motivated by those competing sources of reinforcement
    interfere with each other.

If an organism is receiving reinforcement from one source, how does the
source not providing reinforcement at the moment "compete" with the
source actually providing the reinforcement? If only one behavior at a
time can be performed, how does the one not being performed "interfere"
with the one actually being performed?. If reinforcement is not being
received from a given source, and the behavior associated with that
source is not being performed, is the behavior not being performed still
being "motivated" by the reinforcer not being received?

For lever-pressing to be motivated by a food pellet, the rat must have had
some experience through which the act of lever pressing has become
associated with the consequence of receiving the pellet. Because of this
association (learning), the food pellet motivates lever pressing and not
some other behavior (selective property). The motivation to press the lever
is thus "there" throughout the session, as are any other motivations to do
other things. It is not necessary for the reinforcement from a given source
be "received" continuously in order for that reinforcer to motivate the
behavior with which it has become associated throught prior experience.

If only one behavior at a time can be performed, then two possibilities
present themselves: (a) the behavior with the highest "payoff" may be
performed to the exclusion of the others (exclusive choice); (b) the
behaviors may be alternated in some fashion so that reinforcement from
different sources can be "collected." In the case of (b) it may be possible
for the behaviors to alternate in such a way that that essentially all
opportunities for reinforcement from each source can be taken advantage of.
Alternatively, there may be constraints such that collecting reinforcers
from one source causes missed opportunities for collecting them from the
other source. In the latter case there is competition between reinforcement
sources in the sense that increasing the rate of reward from one source
decreases the rate of reward avaliable from the others. What happens in
these different cases has been the subject of considerable research.

    Now, what determines the reinforcing value of the food pellets,
    that is, the effectiveness of the reinforcer in maintaining the
    behavior that produces it? Again the situation is somewhat
    complex. All else being equal, the rat will respond more
    vigorously for a pellet of food the higher the level of deprivation
    (up to the point where deprivation begins to weaken the animal, a
    region rarely explored in operant studies). So higher deprivation
    leads to the pellet being more highly valued, capable of sustaining
    a higher response rate.

Deprivation, then, multiplies the response-producing power of each
reinforcement by some factor. It is not, however, reinforcing in itself,
because if behavior is not maintained by reinforcement, deprivation
would not be able to affect behavior. On the other hand, it would also
follow that reinforcement by itself is not effective unless there is
deprivation, because of the multiplicative relationship.

Yes.

    The size of the pellet has a similar effect (note: both effects are
    nonlinear. Each successive doubling of pellet size does not double
    the value of the pellet; a log function may be indicated.) The
    flavor of the pellet would also be a factor determining value, and
    several other things as well that I will skip over in the interest
    of brevity.

So physical and biochemical characteristics of the pellet have an effect
similar to that of deprivation. Is this also a multiplicative effect?

Yes and no. Size yes (using perhaps log size rather than absolute size).
Taste enters in, and it is possible to get motivation without any nutritive
value (as in saccharine), so all factors are not strictly multiplicative.

    In addition, there is contiguity between the response and the
    delivery of the pellet to consider when determining value. The
    larger the delay between responding and reinforcement, the weaker
    the effect of the reinforcer on the response (the so-called
    gradient of reinforcement).

Is contiguity simply the reciprocal of delay?

Yes.

How is this delay different from a ratio schedule?

One could program a ratio schedule in which there was a fixed delay between
the completion of the ratio and the delivery of the food pellet. Longer
delays will produce lower rates of responding on the schedule. But delay
may also enter in when considering the effect of the ratio requirement
itself. The first response in the ratio will occur at a delay to
reinforcement that depends on the number of additional responses that need
to be completed and the rate at which those responses are made. If that
rate (the running rate) were constant across ratios, then larger ratios
would impose longer delays between the first response in the ratio and
reinforcer delivery. During the postreinforcement pause (after the pellet
has been consumed) the animal has a choice: return to the lever or engage in
some other behavior (even if that behavior is just to rest). The
probability that the rat would return to the lever would depend on the
operant levels of these competing behaviors, which in turn would depend on
the relative value of immediate reward for the alternative behavior versus
the delayed reward for lever pressing. The prediction would be that larger
ratios should be associated with longer postreinforcement pauses.

If a response occurs, another response
that occurs within the delay time will go unrewarded, so there may be
several responses before the reinforcer occurs. The effect is, on the
average, the same as rewarding only one response out of n, where n is
the average number of responses made during the delay time.

That would depend on whether the rat discriminates the lack of contingency
between response and reinforcement during the delay.

As for the delay imposed by the ratio requirement itself, all responses are
required before reinforcement will be delivered, so this is a different case
from responding during an imposed delay. Once the rat has returned to the
lever, it is easier to continue lever-pressing than to break off and do
something else. Furthermore, once the first response has occurred, a
smaller ratio remains and thus a shorter delay to reinforcement. There is
thus a strong tendency to continue responding once responding on the lever
has begun. The usual view, I think, would be that each response is
reinforced by the eventual delivery of the food pellet, but at different
delays, so that the amount of reinforcement would be slightly greater for
each subsequent response in the series. This would predict that responding
should accelerate during completion of the ratio. This acceleration is
sometimes observed during acquisition, but disappears in the steady state.
A proposed explanation is that with sufficient experience the rat begins to
discriminate the relationship between _rate_ of responding and _rate_ of
reinforcement (as opposed to merely the occurrence of a response, and the
occurrence of reinforcement), and then begins to respond so as to maximize
the rate of reinforcement (but with response rate diminished by the effect
of response costs).

This holds
for interval as well as ratio schedules. Is this reinforcement gradient
used as an explanation for post-reinforcement pauses?

Yes, as described above.

    If there were no other sources of reinforcement in the environment,
    and if lever-pressing had no cost, one might expect the rat to
    press the lever at maximum rate so as to maximize the rate of
    reinforcement, regardless of the level of motivation.

So if we reduced the cost of responding nearly to zero, and eliminated
all alternative sources of reinforcement but the one kind of response,
the animal would be pressing at its maximum possible rate. This suggests
that on a FR1, where the cost per reinforcement is low, the animal would
be pressing faster than on a FR 10, where the cost per reinforcement is
high;

Yes, if that were the only consideration. The problem is that the lower
rates that the higher ratios would sustain would substantially reduce the
overall rate of reinforcement. If the rat were sensitive only to local
contingencies (dependence of individual reinforcer delivery on individual
responses), this would be the clear prediciton. But empirical tests
indicate that rats can learn to discriminate rates of reinforcement as well
as delays to individual reinforcers. Assuming that a higher rate of food
delivery is more reinforcing than a lower rate (to a hungry rat), a
prediction based on overall rates would be that rate would increase so as to
maximize reinforcement rate (but with rate of reinforcement reduced to
somewhat less than maximum in proportion to response cost, which would
reduce the effective value of a given rate of reinforcement).

I know what you're thinking--the theory tries to have it both ways: it can
predict either a direct or inverse relationship here. But the situation is
really no different from the one faced by a PCT analysis of the same
performance. What PCT model you build depends critically on which variable
or variables are being controlled. Logically, several are possible, and
different choices will produce models that behave differently. Whichever
theory you subscribe to, your efforts will be focused on identifying the
variables actually involved and determining how they contribute to the final
performance. And, of course, you will try to develop a model that is
consistent with the data. This criterion will allow you to rule out certain
alternative constructions of the model.

and that if the nutrient value of a pellet were increased, thus
requiring less behavior to produce the same value, the rate of pressing
would also increase.

Well, not the _nutrient_ value per se, but yes, increasing the size, for
example, would be expected to increase the rate, at least under a local
analysis, and assuming that satiation is not approached during the period of
observation.

By the time you get to an analysis based on overall rate of reinforcement
and rate of responding, it is but a short step to a regulatory view. Rate
of reinforcement might, for example, be taken to mean rate of intake, which
would factor in pellet size as well as rate of delivery. Close the loop,
add a set point, and viola, PCT.

The suggestion is also that if there were no (or little) cost and no
competing reinforcers, the organism would maximize the rate of
reinforcement regardless of the level of deprivation (which is the
motivation that gives reinforcers their effect). So even if replete, the
animal would still press as fast as possible.

Well, if the equation is a multiplicative one, at zero motivation, one has
zero output. But under these ideal (unreal) assumed conditions, even an
infinitessimal level of motivation would drive the output to maximum. But
what real system has infinite gain?

So far, here are the general ideas I get from this.

1. Reinforcers play two roles: as "selectors" of behavior and as
"maintainers" of behavior.

Yes.

2. Reinforcers, even when not actually being received, can "compete"
with reinforcers actually being received.

Yes.

3. Behaviors, even when not being performed, can "interfere" with other
behaviors actually being performed.

No. (see discussion earlier in this post) But the occurrence of one
behavior can interfere with (prevent) another behavior, so that it is not
performed, or is performed at a reduced rate compared to the rate if
interference were not present.

4. When reinforcement is maintaining a steady state of behavior,
"motivation" can still vary and thus vary the amount of behavior being
maintained by the reinforcement. So motivation and reinforcement can
vary independently.

Careful--"reinforcement" and "reinforcer" are not equivalent terms. The
reinforcer is not affected by deprivation (it's still the same food pellet).
But its motivating properties relative to the organism being subjected to
deprivation are.

5. Motivation depends negatively on a state of the organism called
"satiation". An approach to satiation reduces motivation.

Yes, if we're talking about some quantity associated with satiation effects.
Not all do: footshock is an example of one that doesn't.

6.Deprivation multiplies the ability of reinforcers to maintain behavior
by altering the "value" of a given reinforcer, as well as increasing
motivation.

Yes, but be careful here. Deprivation might affect the value of different
reinforcers to different degrees (different parameter-values relating
deprivation to value). The motivation to obtain any of these reinforcers
would increase with deprivation, but the motvation to obtain any particular
reinforcer IS its value.

7. Physical characteristics of the reinforcer have effects similar to
those of deprivation: increasing the size of a pellet will have an
effect similar to that of increasing deprivation, increasing the value
of the reinforcement and thus leading to more behavior.

Yes.

Note: the terms put in quotation marks the first time they are used are
undefined explanatory terms.

Well, perhaps I didn't define them in my post, but that is not the same as
their being undefined.

There is at least one puzzling relationship among these purported facts.
If, as you say, each one of them has been established through
experimental observation, this makes them all the more puzzling.

It has been observed that deprivation increases the value of a given
reinforcer, and increasing the value of a reinforcer is observed to
increase the level of behavior maintained by that reinforcer. But the
behavior that is maintained by reinforcement works in the direction of
reducing deprivation, which is obverved to reduce the value of the
reinforcer. At the same time, reducing deprivation decreases motivation
by causing an approach to satiation, which would also work against the
effect of the reinforcement.

Yes, and that is what is observed. However, operant studies in which
satiation might be a problem are designed to minimize this effect by keeping
reward size small and/or session length short.

Motivation is increased by deprivation, so
it is increased by a decrease in obtained reinforcement. But a decrease
in motivation would decrease the effectiveness of a reinforcer in
maintaining behavior, and that would tend to decrease reinforcement and
thus increase motivation.

You are considering only the case in which reinforcement rate is reduced
below that necessary to maintain the current level of deprivaiton. However,
in typical studies a decrease in obtained reinforcement would not increase
deprivation, it would slow the rate at which satiation occurred. The level
of motivation would not increase, it would decline more slowly.

There is a simple explanation of these somewhat contradictory
observations: what motivates behavior (at least after the selection
phase is over) is deprivation, not reinforcement. The more reinforcement
that is obtained, the less the deprivation becomes, and the less
behavior is observed. When the deprivation is completely removed,
behavior ceases. If reward size is increased, the deprivation is removed
faster and behavior ceases sooner. In the limit, a behavior that
produces the entire day's supply of reinforcer in one lump will suffice
to terminate behaving after one response.

A couple of posts ago I suggested what I thought would be the PCT
interpretation of what I had been describing in reinforcement terms. Your
response was to inform me that you knew what PCT would predict, as if I had
been attempting to instruct you on the principles of PCT. I was hoping you
would tell me whether you agreed with my analysis. I now see that you do.
Good!

Maintaining a general deprivation over a 24-hour period does not mean
that short-term deprivations cannot be created and removed. Removal of
short-term deprivation explains why even hungry rats normally eat in
bouts.

I agree.

This is, as you will recognize, the PCT explanation. This approach
focuses not on the range of reinforcements between zero and the obtained
amount, but on the range between the obtained amount and the satiation
level -- in other words, on the error (I am now accepting your
definition of satiation). The reference point is shifted from the
physical zero to the satiation level of obtained reinforcement.

Yes, but let's not forget that we're talking about at least two levels of
control here. The reference for _rate of reinforcement_ (what Motheral's
graphs show) will be zero when the reference for _nutrient level_
(satiation point) has been reached.

That's going to have to be it--this program won't let me enter more than
this in a single post. Anyway, that's probably enough for now.

Regards,

Bruce

[From Bruce Abbott (950705.1815 EST)]

Comment on galley proofs: One of the more subtle forms of torture that
publishers inflict on authors is making them read their own writing--a word
at a time. Looks like time for a break, a chance to relax and catch up on
the mail.

Bill Powers (950704.1650 MDT) --
    Bruce Abbott (950704.1235 EST)

I am unsympathetic with your busy schedule. As I write this, there is a
good rerun of Star Trek going on, and my new electric lawn mower is
waiting to attack the next segment of lawn on the hill west of my house.
In addition, I am very rushed because Mary and I are going to see
_Apollo 13_ in half a hour. We all have to make sacrifices to carry on
the great work.

Yes, Boss.

[Editor's note: had to finish this the next day. Incidentally, _Apollo
13_ is an incredibly good movie -- riveting!]

I saw it too: outstanding! The making of the movie was shown on cable.
Among the consultants were some of those who worked in the control room at
the Houston Space Center. They said that after they had spent a few days on
the set, the control room was so accurately recreated that they found
themselves looking for the elevator when they left the room. The _real_
control room had been on the third floor.

    I don't see why average motivation would be expected to _increase_
    with ratio size; in fact I would expect the opposite, due to the
    increasing response cost.

Remember that as we move leftward on the Motherall curve, we are moving
in the direction of reduced reinforcement. Since I am defining
motivation as the difference between steady-state obtained reinforcement
and the satiation level of reinforcement (to which I believe you
agreed), this means that moving left on the curve is moving in the
direction of increased motivation. As ratio size is increased, the
observed points _do_ move left on the Motherall curve and motivation, as
defined, increases. So what motivation is _expected_ to do is
irrelevant, except for purposes of invalidating any model that led to a
different expectation under the same definition.

Looks like I need to clarify this satiation concept. I'm viewing the
control system involved as comprising two levels. The top level involves
something like nutrient level and/or stomach loading; let's just call it
"hunger" perception. This perception varies from less than zero to some
maximum positive value. Zero represents the absence of hunger and the
beginning of satiation (as satiation increases the values go negative); it
is also the presumed set point of the hunger system. [Note: This is a
considerable simplification of the real system/systems involved.]

Food deprivation increases the hunger level. As hunger rises above its
reference level this produces a rising error signal, which represents
"hunger motivation." This signal serves as the reference for the
lower-level system, which sets the desired rate of food consumption (which
is controlled via rate of lever pressing on the ratio schedule).

If rate of food consumption is low enough relative to the rate at which the
food is "burned," deprivation level will increase over time and hunger
perception will increase toward its maximum (if not already there). At
typical schedule parameters, however, the rate of food consumption is large
enough to reduce deprivation level over time. Thus lever-pressing serves to
reduce error in the "hunger" system (but not rapidly) over the course of a
session. Whatever changes have occurred during the session will be
corrected by the experimenter after the session by adjusting the amount of
supplemental feeding.

In this construction the rate of lever pressing will depend on the level of
hunger motivation (proportional to deprivation level if the hunger setpoint
remains at zero) and on the amount of reinforcer delivered per lever-press.
The amount of reinforcer delivered per lever-press would depend inversely on
the schedule ratio and directly on the size of the food pellet, and perhaps
on the nutrient content of the pellet.

If food is consumed faster than it is expended, hunger motivation declines
over time. If this continues long enough, the satiation point is reached
(hunger perception reaches its reference level), the error signal in the
hunger system goes to zero, and thus the reference for the reinforcement
rate system goes to zero. As the reference declines, the rate of
lever-pressing declines to keep the rate of food intake close to its
reference level. At satiation the reference for food intake is zero and
there is no lever-pressing.

The level of hunger (proportional to deprivation level) thus determines the
rate of food consumption. If we follow the right limb of the Motheral curve
down to the abscissa, that value is the rate of food intake that would occur
if the food were freely available and thus is a measure of the level of
hunger. If before I claimed that this would represent the satiation point
then I wasn't thinking clearly on the matter. The satiation point is the
reference level of the hunger system.

What confused me, I think, was thinking that at satiation the rate of
lever-pressing would be zero. On the graph the zero rate of lever-pressing
corresponds to the set-point of the reinforcement-rate system. But it's the
wrong system. Satiation is the set-point of the hunger system, not of the
reinforcement-rate system.

The effect of varying the level of deprivation (hunger) is to alter the
reference level for rate of food intake. With a constant pellet, the rate
of food intake is controlled by varying the rate of lever-pressing. A lower
level of deprivation should therefore reduce the rate of lever-pressing at
all ratio requirements. Susan Motheral's two curves at 98% and 85%
deprivation confirm this effect.

   Furthermore, one would expect a shift to other behaviors as the
    motivation for lever-pressing _decreased_, not increased.

We are defining motivation as satiation minus obtained reinforcement.
This means that as motivation increases, reinforcement must be falling.
It is the decrease in reinforcement that leads to an increased error
that leads to increased behavior and, if large enough, results in a
shift to other behaviors, either to established other behaviors that
could produce greater reinforcement or to a search pattern. The added
effort involved only contributes to the general increase in error.

Now that I've pulled the rug out from under you by reidentifying (correctly
this time, I think) what satiation corresponds to, this no longer holds.
The relationship involving lever-pressing and reinforcement rate still
holds, but the interpretation changes.

What does happen from a reinforcement point of view is that as the ratio
requirement increases the net reinforcement (benefit - cost) maintaining
behavior on a given ratio schedule decreases. At some point the net
reinforcement reaches zero and the response extinguishes.

    Rather than my presenting another verbal description, perhaps it
    would be best to wait until I can develop the mathematical model
    that will express the sort of relationships I have in mind.

The relationships I describe are present in the mathematical model in my
last post.

Yes, I know. I'm talking about my model, not yours.

    I've said this before, but it bears repeating. I do not agree with
    this analysis. Most operant studies using ratio schedules do not
    employ conditions represented by the left side of Motheral's curve.

How do you know? To know where you are on the curve, it is necessary to
explore the entire range of schedules, and also to know how the
conditions you have chosen (particularly reward size) affect the
position of the peak of the Motherall curve. Without a model to allow
you to predict where the peak will be, there is no way to say _a priori_
which side of the peak you are on. If you adjust the reward size to
avoid a leveling off of an expected positive relationship between
reinforcement and behavior, you are automatically adjusting the
reinforcement value to assure operation on the left side of the peak. I
doubt whether very many (if any) experimenters have been aware of this
peak-shift phenomenon, especially considering that the Motherall curves
have seldom been commented on as puzzling (to my knowledge). Did Skinner
ever mention such curves as a paradoxical effect? In most studies of
which I have heard, behavior rates in a single task are observed to rise
with reinforcement rates.

This may come as a surprise, but the amount of basic research Skinner
published is rather miniscule. Most of the early results in this field were
described mostly in qualitative terms (e.g., patterns of behavior that
develop on a given schedule; whether one schedule supports higher rates of
responding than another when reinforcement rates are equated, etc.) Almost
no work was done with ratio schedules, and most of this did not focus on the
effect of varying the ratio requirement.

On interval schedules, shorter intervals (higher reinforcement rates) do
indeed produce higher rates of responding (up to an asymptote). On ratio or
interval schedules, larger rewards (equivalent to higher reinforcement
rates) produce higher rates of responding. Perhaps manipulating the ratio
was not seen as a "proper" way to manipulate reinforcement rate, since the
ratio imposes a strict relationship between rate of reinforcement and rate
of response (no degrees of freedom).

    Overly large ratios produce "ratio strain," a breakdown of
    "schedule control." Responding on the ratio becomes unstable and
    may cease altogether. _This_, I believe, is the region represented
    by the left limb of the curve.

I agree that if you make the ratio high enough, this phenomenon may well
be seen. But it is perfectly possible to operate in a stable manner on
the left side of the peak. If my model is correct, all that the left
side of the peak represents is a region where, as reinforcement
decreases, the output gain is decreasing faster than the error is
rising. There is still gain, and it is negative: a line drawn from the
satiation point (the reference level) to the leftmost point on the
Motherall curve still has a negative slope.

Well, it's a nice model, but it isn't consistent with my intuition. And I
still hold that this region is not where operant studies are typically run.
I'm not sure what evidence I need to present to support this.

I agree with you that this effect is likely to be a result of a
nonlinear benefit-minus-cost function. It says that we are working in a
region where a decrease in behavior decreases the benefit faster than
the cost decreases (a comparison of slopes). But in this region, the
benefit still exceeds the cost (a comparison of magnitudes): it is still
better to behave than not to behave.

Yes, and it may be better to spend less time doing it, and more time doing
something else more "rewarding" in other ways.

It might be that what you say is right: that in many experiments, it is
observed that lower ratios go with higher reinforcement rates and lower
behavior rates. But if that is the predominant observation, how did
psychologists ever get the idea that behavior is _increased_ by
reinforcement? The only case in which this idea applies is when the
organism is shifting from one kind of behavior to a different kind. Once
a behavior has been selected, what now is the relationship between
amount of reinforcement and amount of behavior? Nothing in my reading
has suggested that behaviorists think it is any different.

See above discussion on reinforcer amount and on interval schedule size.

Re: Rachlin's book

Rachlin's book was written at just about the lowest level possible for an
audience assumed to have little or no quantitative skill. In many ways
Rachlin oversimplifies in order to present the basics in a straightforward
way. It is not surprising that the issues you raise are given little or no
attention in the book. That said, there are problems with many of these
analyses you could drive a truck through. What do you think got me started
looking for some other framework?

To be fair, the paradoxes you note have been noted by workers in the field
and there have been theoretical efforts made to deal with them. For
example, the "partial reinforcement extinction effect" (PREE) had been dealt
with by several theorists. One view brings in the notion of discrimination:
That the rat working on a partial schedule of reinforcement learns that
reinforcement occurs infrequently during responding. The less often a
response is reinforced, the more difficult it becomes to discriminate the
switch to extinction, in which no responses are reinforced. (I've given you
the verbal explanation; there are quantitative models of this.)

    Jim, you've got it half right. Just substitute PCT for this silly
    notion of "response deprivation." (;->

The "response deprivation" idea confuses two levels of control. Staddon
does the same thing, by eliminating the reinforcer and substituting the
behavior of ingesting the reinforcer -- it is eating behavior that
reinforces bar-pressing behavior. Both of these approaches lose track of
the controlled variable, and confuse the controlled variable with the
means of affecting it. This may be because the real controlled variable
is inside the organism and must be deduced -- out of bounds for a strict
behaviorist. Obtaining and ingesting little tan or green objects is not
reinforcing; what is reinforcing is the effect on the organism of doing
this. So we have two behaviors, bar-pressing and eating, which when
performed in a fixed alternation can control something in the nervous
system of the organism that is affected by this process.

Now that's just what _I_ was thinking. You must have peeked. (;->

Regards,

Bruce

[From Bruce Abbott (950706.1730 EST)]

Bill Powers (950706.0905 MDT) --

I'm running off to class in 20 minutes and won't have time to reply later
tonight, so here's just a few quick comments on this just-arrived post.

So far so good, but I would make some small modifications. The argument
is simplified if we just identify the state of hunger as being
proportional to the deprivation -- i.e, to the error signal. So the only
reference signal needed is for the stomach loading or whatever -- a non-
zero reference signal specifying a non-zero level of the controlled
variable. The relationship of the consciously-experienced state we call
hunger to the variables in these low-level control systems is not
necessarily straightforward -- the subjective experience of hunger may
be associated with sensed efforts in lower-level systems to do something
about error signals (such as churning in the stomach, etc.).

Fine with me.

So this gives us two levels of satiation: one associated with control of
rate of ingestion, and the second associated with control of stomach
loading. I would like to use a uniform definition of satiation that is
consistent with our operational definition of reference level. The
reference level of a variable is that level of the variable at which
behavior just goes to zero. The degree of deprivation, which is the
error signal, is then just the reciprocal of the degree of satiation.
When error is exactly zero, we have infinite satiation.

I had thought about this when considering the redefinition of satiation as
involving only the higher-level system, that if satiation is reached when
hunger level reaches zero (reference level), then it might make sense to
talk about "satiation" of rate of food intake and the like. However, I am
actually thinking of satiation as just the negative side of a continuous
scale (satiation is negative hunger) rather than a reference point. If the
reference is zero, then as hunger decreases, the reference is where
satiation begins. This definition avoids the business of infinite
satiation. If we use stomach loading as the controlled variable, there is
some value of stomach loading (NOT zero) at which satiation is reached;
adding more pushes us into the satiation region. Thus both hunger and
satiation are names for error, but on opposite sides of the reference level.

I'll have to pass over the rest for now...

Bruce, of course, is hopelessly contaminated by PCT, so when he starts
pushing the elements of reinforcement theory around to get them into
some kind of consistent form, he's inevitably introducing PCT concepts.
By the time we have finished, his reinforcement model will probably be
totally unacceptable to his colleagues in EAB :slight_smile:

You're probably right about the contamination. It the resulting model is
totally unacceptable to my colleagues in EAB, I'll be ecstatic.

He will find himself
ostracized by his former friends, and he will have to join the ranks of
the institutionally homeless like you and Tom Bourbon and all the others
who have contracted this disease.

I'm not too worried about that. For one thing, nobody is going to throw me
out of IPFW for teaching PCT or doing PCT research. They're having too much
trouble finding PhD's willing to work here for this wage... (:->

Regards,

Bruce

[From Bruce Abbott (950707.2125 EST)]

Bill Powers (950706.0905 MDT) --
    Bruce Abbott (950705.1815 EST)

    What does happen from a reinforcement point of view is that as the
    ratio requirement increases the net reinforcement (benefit - cost)
    maintaining behavior on a given ratio schedule decreases. At some
    point the net reinforcement reaches zero and the response
    extinguishes.

But think about this. As the ratio requirement increases, the net
reinforcement decreases, and at some point the response extinguishes.
But what we observe in the Motheral curves is that as the ratio
increases and the reinforcement decreases, the behavior rate
_increases_. We observe the decrease in reinforcement, and can
reasonably deduce that the _net_ reinforcement is decreasing even faster
than what we observe because of the increased costs, but we do not
observe a decrease in the behavior rate. Instead, we observe an
increase. Only when the schedule ratio exceeds a certain amount, and the
reinforcement rate falls below a certain small percentage of the
reference level, do we start to see a decline in the behavior rate (and
eventually, we presume, extinction). What you are talking about is only
the region to the left of peak in the Motheral curve. To the right of
the peak, we do NOT observe the relationship you describe above.

Yes, exactly so. To quote the King of Siam in _The King and I_, "Tis a
puzzlement!" (if you are a reinforcement theorist). The net-value analysis
I described "explains" why responding declines with increased ratio (left
limb of the curve: motivation is declining. The same explanation applied to
the right side of the curve clearly predicts the same trend--declining
responding with increasing ratio--which of course is contrary to observation.

Without getting into the details at this point, a variety of theoretical
solutions have been proposed. Some of them are even billed as "regulatory"
models, but all those I am aware of miss the mark at some point or other, in
my opinion (e.g., Timberlake's "behavior regulation" model; Allison's
"response deprivation" model).

    On interval schedules, shorter intervals (higher reinforcement
    rates) do indeed produce higher rates of responding (up to an
    asymptote).

Wait a minute. I can see that higher rates of responding would produce
higher rates of reinforcement (up to an asymptote); that is merely the
nature of the feedback function in an interval schedule.

Yes, higher rates of responding _on a given schedule_ will produce higher
rates of reinforcement, up to an asymptote. But that's not what I'm talking
about. Interval schedules tend to produce far more responses than required
to collect each reinforcer. With the levels of deprivation and size of
reward generally used in these studies, we're operating in essentially the
flat (asymptotic) region of the curve relating rate of responding to rate of
reinforcement. As we increase the size of the interval on the interval
schedule, response rates decline. Or, to put it another way, as asymptotic
rate of reinforcement on the schedule is increased, rate of responding goes up.

But are there
the equivalents of the Motheral data for interval schedules where
obtained reinforcement rates are plotted against behavior rates for a
wide range of schedules? I very strongly suspect that we would see the
same general kinds of curves as for ratio schedules, although the shapes
might be different due to the nonlinearity in the schedule function.
After all, if each successful behavior produced a generous amount of
reinforcer, the animal would probably not press at a high rate, unless
the schedule were set so that reiforcements came only rarely.

I assume so but I'll have to check--when I get time.

RE: region of Motheral curve left of peak:

    Well, it's a nice model, but it isn't consistent with my intuition.
    And I still hold that this region is not where operant studies are
    typically run. I'm not sure what evidence I need to present to
    support this.

The required evidence would involve data that not only reproduce the
Motheral curves, but do so over a wide range of reinforcement sizes.
There is no absolute scale of reinforcement size; about the only
comparison that can be made is in terms of amount received in one
reinforcement against amount consumed while free-feeding, both per unit
time. Reinforcement sizes are probably adjusted to produce reasonable-
looking data, and what is reasonable depends on what you believe should
be seen. If you think it reasonable that more frequent reinforcement
should produce more behavior, you will adjust reinforcement size until
that is what you see.

I don't think any adjusting has been done to force data to match theory.
Pellet size has been selected to sustain reasonably high levels of
responding while not satiating the animal too rapidly (one wants to be able
to observe behavior long enough to see steady-state responding develop).
I'd bet good money that conditions generally studied fall on the right limb
of the Motheral curve (on ratio schedules). On other schedules where rate
of reinforcement is not so rigidly linked to rate of responding (e.g.,
interval schedules), behavior rate tends to increase with reinforcement rate.

So we have two behaviors, bar-pressing and eating, which when
performed in a fixed alternation can control something in the nervou
system of the organism that is affected by this process.

    Now that's just what _I_ was thinking. You must have peeked. (;->

That is a remark that bodes well for our project.

Glad to hear it. I gather you have some acquaintance with Allison's model?

Regards,

Bruce