Fixed Ratio Data, etc.

[From Bruce Abbott (950714.1215 EST)]

Bill Powers (950712.0825 MDT) --
    Bruce Abbott (950711.1615 EST)

    In learning one would expect that following a response with a so-
    called reinforcing consequence would "strengthen" the behavior
    (relative to other behaviors not so reinforced).

Right, so at least the relationship between reinforcement and behavior
goes in the "right" direction while a given behavior is being selected
from among alternate behaviors. Any mechanism that gradually shifts from
a search or trial-and-error pattern to a consistent pattern that
produces a specific consequence would create an appearance that supports
the basic reinforcement model (or would not contradict it).

Yes. The problems develop when the same principles are applied to account
for steady-state behavior. This latter situation is one in which I have
always felt that a "regulatory" model (as I used to call control-system
models) provided a far simpler account of the observations.

    Maintained performance is another game altogether--a dynamic
    equilibrium involving all sorts of effects. You don't necessarily
    expect _rate_ of responding to be directly related to _rate_ of
    reinforcement under these conditons. There are too many other
    factors to consider and could alter this simple relationship.

No, I can't accept this explanation. I hope you're just playing Devil's
Advocate. These are just words that vaguely indicate a lot of possible
interfering factors without naming any of them or laying out what their
individual effects would be -- or showing that they exist. This is an
age-old excuse that is used when a prediction fails: well, the spell
really does work, but the Moon was rising and Og was probably casting a
counter-spell while we were all asleep, and anyway you probably didn't
believe hard enough. That's why you still have warts.

I'm just suggesting the general "approach" used by reinforcement theorists,
as there are several more specific solutions which have been proposed. When
prediction fails, it is only natural to try to account for the unexpected
results by attempting to construct a model based on known principles. But,
contrary to the implication of your paragraph, these models are themselves
submitted to test. If you look at any issue of JEAB, for example, you will
find studies being conducted to evaluate alternative models of behavior as
observed in some specific situation, such as multiple schedules or
matching-to-sample.

We will have the opportunity to evaluate some of these specific alternatives
against specific PCT models when be begin to collect our own data. At the
moment I'm still trying to "get current" as to what those specific views
actually are, and what they would predict. If I'm being vague in MY
description, it's because I don't yet know enough about them.

Getting back to the issue of "devil's advocacy," yes, I'm just trying to
convey a general notion of how reinforcement theorists would approach the
problem posed by data such as those relating rate of responding to rate of
reinforcement on various schedules. Staddon, Timberlake, Allison, Mazur,
and others have proposed specific approaches, and there is currently much
debate about the merits of each relative to the others as well as empirical
research designed to test implications of these models. I'm not asking you
to "buy" any of these solutions; I'm asking you to accept the idea that
there is a general awareness of the problems posed for traditional theory by
data such as these and that research is underway in various labs to evaluate
alternative solutions, as opposed to simply offering glib verbal
explanations that make the anomalies seem to go away.

    I'd like to see your model extended so that kt emerges from the
    competition of the two systems rather than being fit to the data.
    But perhaps that can wait.

That's worth trying. I found, by the way, that the left limb of the
Motheral curve can be created (contrary to what I said earlier) by using
a zero offset and a cost function that simply goes as the square of the
error. You were right about that. And now that I think of it, the cost
should really go as the square of the behavior rate, not the error --
I'll try that when I'm finished here.

Good, I'm anxious to hear of the result.

    The reason these rats continued to respond on the ratio schedule at
    ratios up to 5000 is that this was their only way to obtain food--
    it's that or starve. Food deprived rats working on a ratio
    schedule in a 1-hr session will give up responding on ratio
    schedules far less demanding. It's interesting to think about why.

It is obviously necessary to provide enough food or water at the
extremes of the schedule to keep the animals alive when they are in the
apparatus full-time. This puts a lower bound on the reinforcer size. I
have speculated that this is the size that determines where the Motheral
peak will be found. The Timberlake data show that as the ratio
increases, the amount of reinforcer actually obtained per access
increases by a factor of almost 10, so the reinforcer size is not fixed.
This is a problem with using "access" as a reinforcer. You would really
need a two-level control model for such a case, with the lower system
controlling rate of ingestion, and the reference level for rate of
ingestion being determined by error in the obtained amount of reinforcer
per unit time.

The simple notion of the cost-benefit function also suggests that the curve
should move to the right with larger reinforcer amounts, as I noted in an
earlier post. In the data cited by Timberlake, the authors were taking a
regulatory view of feeding and drinking and so offered food (or water) under
nondeprivation conditions within what was essentially the home cage, 24
hrs/day. In one experiment they could earn access to the food or water by
lever-pressing and could then eat/drink as much as they desired in that bout
of access. Failure to access the food/water for 10 minutes ended the bout,
after which the rats had to respond on the lever again to gain access.

The authors suggested that this situation represents a laboratory analog to
what the rats would have to do in the wild. Lever-pressing for access is
equivalent to "foraging" for food or water; once a source was obtained the
rat would then consume a "meal" or "drink" of a size of its own choosing.
The authors suggested that the pattern of eating/drinking the rat would
prefer might be influenced by such considerations as the effort required to
obtain the food/drink, environmental dangers, and so on (I will omit the
details here). Under the usual conditions of the rat's natural environment,
it might prefer to consume small meals frequently. However, if foraging is
costly or dangerous, these conditions might lead to the rat's eating larger
meals less frequently, the larger meals tending to compensate for the reduce
frequency. The study was designed specifically to evaluate this hypothesis,
and that is why meal size or amount drunk per access was placed under the
rat's control.

The authors' analysis would suggest that meal size or amount drunk (per
bout) and access frequency are both controlled perceptions in the service of
a higher-level system controlling overall consumption. Constraint of one
system (e.g., meal frequency) leads to error in the higher system which
alters the reference level for the other lower-level system (meal size;
drinking bout size) so as to correct that error; the result is a
"compensatory" increase in meal size or drink size).

Remember that as intervals _increase_, rate of reinforcement is
_decreasing_. Your guess applies to the left side of the Motheral curve;
as interval increases toward infinity, rate of both reinforcement and
responding decreases toward zero. A decrease in the interval corresponds
to moving rightward on the Motheral curve.

O.K., so you're suggesting that on interval schedules, results such as those
reported by Catania and Reynolds indicate that at the smallest intervals
tested, the pigeons were already operating on the left side of the curve. I
can follow your reasoning easily enough, but there's something about this
that's bothering me that I haven't been able to put my finger on yet. It
seems somehow related to the use of deprivation in these experiments; I'll
have to think about this a bit more to see whether there's anything to it.

Rick Marken (950711.1615)

    So reinforcement stops "strengthening" behavior when responses have
    been learned and are being "maintained". How does the reinforcement
    know when to stop strengthening and start maintaining? What are all
    the sorts of effects that maintain dynamic equilibrium (betweem
    responses and reinforcemnts, I presume)? What factors enter the
    picture during the "maintaining" stage that alter the simple
    relationship between reinforcmement and response rate that
    presumably existed during the learning stage?

Good questions. It's obvious to me that two quite different mechanisms
are involved; calling them both "reinforcement" is a mistake. During the
"selection" phase, one process is at work that shifts the proportion of
time spend behaving in different ways and places, the shift slowing down
when a particular way or place produces reinforcement. This process
continues to work even after the selection is complete, but keeps
converging back to the same selection as long as the same amount of
reinforcer is obtained. Some version of the E. coli effect might serve
as a model of this process, or perhaps there is a systematic method that
would also work.

The other process is simply control. When control is predominant, we
have the negative relationship between behavior and reinforcement, with
a reference level at some (size times reinforcement rate). After the
right behavior is found, the loop gain of the control system gradually
increases until there is a comfortable margin of benefit over cost of
behaving.

Anyway, that's the general model I'm anticipating now.

Yes, that's the way I see it, too. In attempting to use one concept
(reinforcement) to explain both acquisition and maintenance, reinforcement
theorists have confused two distinct processes which must be treated separately.

Bruce Abbott (950711.2120 EST)

    Eventually, Skinner asserted, you will be led back to observable
    events in the environment.

Yes. It's hard to find explicit statements of his underlying model, but
occasionally they show up. In _Science and human behavior_ (p. 28) we
find this:

    Eventually, a science of the nervous system based on direct
    observation rather than inference will describe the neural states
    and events which immediately precede instances of behavior. We
    shall know the precise neurological conditions which precede, say,
    the response "No, thank you." These events in turn will be found to
    be preceded by other neurological events, and these in turn by
    others. This series will lead back to events outside the nervous
    system and, eventually, outside the organism.

This is clearly the basic stimulus-response model. Neural inputs cause
neural outputs. I've never understood how EAB types could deny that they
are S-R theorists, when this model underlies all their thinking. I
suppose they must identify "S-R theory" with a _particular version_ of
this model, and because they use a more complex version, think that they
are doing something different. In fact, S-R theory utterly dominates
most branches of scientific psychology, including EAB.

The problem with this is that you are using the term "S-R" to refer to all
systems of lineal causality. This is guarenteed to confuse behaviorists,
especially those in EAB, and leads to your being accused of being about 50
years behind the times. S-R theory as used in psychology refers to a
reinforcement theory (especially the one developed by Clark Hull and Kenneth
Spence) which assumed that all behavior is essentially a series of reflexive
responses to eliciting stimuli. A specific stimulus was supposed to produce
a specific response just as in the doctor's knee-jerk reflex. This doctrine
was repudiated long ago, although lineal causality is, as you note, still
very much with us.

Bill Powers (950712.1730 MDT) [responding to Oded Maler (950712)] --

In both cases we end up with an equilibrium between opposed forces. But
in the first case, the equilibrium is passive, while in the second it is
active.

In Rick's earlier post in which he first responded to my suggestion that
reinforcement theorists would handle the ratio data using an equilibrium
model, he noted that equilibrium models and control models are different.
In my reply, I was tempted to point out to Rick that control models ARE
equilibrium models, but after further thought decided that I'd just be
opening another can of worms. The critical difference, as you note above,
is that the control system uses its own source of energy to _actively_
oppose the disturbance (i.e., gain > 1).

Models such as Allison's response deprivation model are equilibrium models
but do not control because they do not include this active opposition
element. It should be relatively easy to show which is correct (equilibrium
versus control) by determining whether there is active resistance to
disturbances.

Regards,

Bruce