Modeling Operant Conditioning

[From Bruce Abbott (950301.2040 EST)]

Bill Powers (?: no time/date stamp)

I've spent considerable time thinking about your diagram and the problem
you're trying to solve. I realize once again why I have done so little
modeling with reorganization. When you actually start figuring out what
has to be modeled in the situation you describe, the problems simply
multiply.

Was there a post to which this is a reply? What are you talking about? I
feel as though I've come in in the middle of an ongoing conversation.

One thing that helps is to put yourself inside the rat. Mary and I did
that for a while this morning. She moved her finger around and whenever
she made a move in the right direction, I said "click," meaning that
water was now available. At first I said click for any move vaguely in
the right direction. Then I demanded more and more, and eventually got
her finger to touch the cup I had in mind as the "instrumental
behavior." Then each time she touched the cup I said "click." She would
run her finger over to where we sort of agreed the water was, then back
to the cup to touch it again. Then I changed the "schedule" so two
touches were required. She touched the cup -- no click! So immediately
she touched it again and I said "Click" because that was two touches.
About two trials later she was touching the cup twice in a row and
getting a "click" right away.

This little game reminds me of the old game of "hot and cold." Move this
way and everyone says "you're getting hotter," move that way and they say
"no, cold, very cold!" The goal, of course, is to find whatever is hidden
and if the (negative) feedback is good, you inevitably will. One property
of "reinforcement" (and "punishment" as well).

For the magazine-trained rat, the click of the feeder is a signal that food
has arrived in the food cup; so control passes immediately to the
position-sensing perceptual system with reference set to the cup's
coordinates. But something else has happened as well, unobserved. Upon
devouring the food, the rat will probably return to the location it occupied
when the click occurred and more-or-less repeat whatever it was doing at
that time. We say that the rat is learning "what to do," but in reality it
is learning what to perceive, and in what order.

The desired repetition may not occur initially; it is difficult for the rat
to know which aspects of those preceptions active at the time of the click
should be associated with the click, a difficulty known as the
assignment-of-credit problem. Was it the perception of pressing down on the
lever? Being in this corner? Looking at the cup? Repetition helps to sort
it all out: only those perceptions consistently present at the time of the
click (or at some time shortly before then) remain likely candidates. In
some way the rat must identify those perceptions that correlate best with
the click, even if at some temporal remove; as you put it, some rather
complex process going on in the rat.

But for this to work, there must already be a lot
of organized behavior available, and there's no real reason to think
that simple random reorganization is the only operative factor.

Exactly. E-coli may have no other option, but rats can draw on a quite
varied repertoire of "solutions" to try.

For example, after a "click," Mary would run her finger over to where
the water was agreed to be, and then _back to where the finger was when
the click occurred_. Do real rats do that? After the interruption of
going to get the water or whatever, do they try to re-establish the
conditions that held at the time the water was seen to be available, and
continue the search from there?

Yes, they often do. However, other perceptual systems may have control--the
smells, the cracks and crevices in the corner of the chamber, the view of
things lying out of reach in the pan below the grid floor--and control may
return to these systems after the food has been eaten. Only after several
repeated coincidences between perceptual state and click will control be
asserted in such a was as to more or less re-establish that state, at least
when the rat is FIRST trained. But the process, as you surmise, is
definitely not random.

I'm just now starting to think again about these issues, trying to fit what
I observe to be happening in the operant chamber into a PCT model. The
ramifications are enormous. I believe that a proper PCT-based model will
account for the effects we naive behaviorists refer to as reinforcement and
punishment, effectively explain so-called conditioned reinforcement and
punishment, and show why the term "self-reinforcement" is absurd. (And
that's just for starters.) But there are still a lot of gaps to fill,
particularly how to deal with what in operant terminology is called
"stimulus control." What mechanism selects which control system will have
current access to the lower-level systems required to produce the necessary
behavioral outputs? On what basis? How to certain perceptions (e.g., the
cue light is on) determine which goals will be pursued and/or which systems
will be engaged to pursue them? I could go on, but speaking of food, it's
8:40 pm and I haven't yet had my dinner. I'm going to go looking for a nice
dose of those "dormative principles" that make consuming certain substances
such a pleasant goal...

Anyway, all these ramblings were triggered by your post, Bill. What was the
question to which you were responding?

Regards,

Bruce

(From Bruce Abbott [950302.1600 EST])

From Rick Marken (950302.0930)

That sounds very interesting! What does PCT have to offer EAB, Bruce? So far,
it has apparently been an offer they _can_ refuse :wink:

I dunno--I'm hoping to have that figured out by March 23rd! Seriously, I'm
planning to introduce PCT, list its similarities to EAB (single-subject
approach, etc.), point out at least one major difference (behavior: control
of perception not stimulus: control of behavior), illustrate the basic
control-system model, then apply the model in the context of an asthma
self-management program designed to teach asthmatics how to control their
asthma symptoms. I hope to show that PCT provides a better model of the
self-management process than does one based on Skinnerian principles.

Bruce Abbott (950301.2040 EST) --

This little game reminds me of the old game of "hot and cold." Move this
way and everyone says "you're getting hotter," move that way and they say
"no, cold, very cold!" The goal, of course, is to find whatever is hidden
and if the (negative) feedback is good, you inevitably will. One property
of "reinforcement" (and "punishment" as well).

Are you saying that "good" negative feedback is a "property" of
"reinforcement" and "punishment"?

Wrong connotation. By "good" I meant to convey "accurate." These words
will not allow the goal to be reached if they are untrue. And what I'm
saying is that reinforcers are behavioral consequences that alter controlled
perceptions; these consequences are in the feedback loop between behavioral
output and perceptual input.

Saying that "control passes immediately to the position-sensing perceptual
system" implies that the rat was not controlling sensed position before the
feeder click.

Given that the rat must be standing at the lever to press it, you are no
doubt correct that position is being controlled throughout, but in other
situations it might not be. I agree with you that here what probably
changes is simply the reference position. On the other hand, if the rat had
merely to rear back in order to make the feeder click, it is quite possible
that the rat would not resist being moved around prior to rearing. If so,
how would you account for position being controlled at one moment, not
controlled at another? The answer, of course, is that this is determined by
another, higher-level control system. I don't think I said anything
contradictory to that.

I believe that a proper PCT-based model will account for the effects we
naive behaviorists refer to as reinforcement and punishment,

We already know it does.

Granted. I was just trying to be complete.

effectively explain so-called conditioned reinforcement and punishment,

Yes. I'm sure it can but I don't know of any working PCT models of this
phenomenon yet. It probably involves learning to control a variable (such as
the relationship between click and position relative to bar) in order to
control another variable (such as the amount of time spent pressing the bar
to get food).

Just what I had in mind, although there is probably more to it than that.
Conditioned reinforcers may simply mark progress toward the primary goal.
Or hey may indicate or determine whether control of some perception is
possible, (e.g., if I have money, I can buy things I want.) But I don't
want to rule out the power or raw association just yet (drink Whizzo because
the ad associates Whizzo with sports cars and pretty girls).

and show why the term "self-reinforcement" is absurd.

Actually, it doesn't sound that absurd to me. If reinforcement is the
reference state of a controlled variable, then the state of that variable
(such as rate of food input) is being controlled in order to produce a result
(reference rate of food input) that the organism wants for _itself_. The term
"self-reinforcement" may be a bit vague, but I don't think it's absurd.

The term "self-reinforcement" is usually applied in the context of trying to
establish control over some perception, such as one's weight. "If I stick
to my diet today, I'll permit myself to take in that movie tonight that I've
been dying to see." According to the traditional view, you stick to your
diet that day because sticking to the diet is being reinforced through the
contingency between eating and movie. The problem with this is that you
could at any moment elect to see the movie, whether you stick to your diet
or not. This decision to attend the movie would, according to theory, be
reinforced by the enjoyment attending the watching of the movie. What's to
stop you? Nothing but your desire to stick to your diet. So, is the movie
contingency responsible for your sticking to your diet, or is your diet-goal
responsible for your (a) setting up the contingency in the first place and
(b) sticking to it, when you could easily circumvent the contingency?

Note how different the situation is if some outside agency determines access
to the movie. If you want to see the movie, you have no choice but to stick
to your diet. If the movie serves as a sufficiently effective reward, you
will stick to your diet in order to see the movie. But if YOU determine
whether or not to go to the movie, can it be said that access to the movie
is motivating your sticking to the diet? Hardly. If you really wanted to
see it, after cheating on your diet you'd go anyway. Self-reinforcement is
thus self-contradictory.

But there are still a lot of gaps to fill, particularly how to deal with
what in operant terminology is called "stimulus control."

Now _there_ is a term that is absurd. It refers to an impossible entity -- a
"stimulus that controls". PCT deals with this kind of "stimulus control"
very easily; it says that it doesn't exist. What does exist (and what people
are probably describing when they use the term "stimulus control") is
"response to a disturbance of a controlled variable".

A pigeon is trained to peck at a response key (a plastic disk on the wall)
when the key is illuminated in red but to turn in a complete circle when the
key is illuminated in green. Either way the response, when completed,
produces brief access to a hopper full of grain. We operant guys would say
that the two responses, keypecking and turning, are both under "stimulus
control," meaning that the probability of observing each response depends on
which stimulus (red or green) is currently present. What controlled
variable is being disturbed? Why does behavior change when the
"discriminative stimulus" changes? Does PCT really say that this corelation
between observed behavior and keycolor does not exist? If it does, I'm
afraid that PCT is contrary to fact. Rick, you can't make something go away
by asserting that it does not exist. [Unless you've become a solipist!]

A more productive approach would be to offer an explanation for it from
within the PCT framework. By asserting that stimulus control is "response
to a disturbance to a controlled variable," you are only stating the basic
PCT assumption that all behavior is error-driven. That is not at issue.
What is at issue is to clearly explain what perceptions are disturbed and
how those disturbances trigger a shift in what is being controlled at a
lower level and/or how. How many levels are involved? At what level does
the change in discriminative stimulus act? At the program level? Some other?

I didn't say that it would be particularly difficult to develop a coherent
PCT account of these phenomena, only that it needs to be done and that I'm
starting to think about it. I'd certainly welcome your insights.

What mechanism selects which control system will have current access to the
lower-level systems required to produce the necessary behavioral outputs?

Why do you think that such a mechanism is necessary? If you run my
spreadsheet model you will see that the same set of lower level control
systems can be used _simultaneously_ as the means of controlling several
different higher level variables.

Try juggling three beachballs while doing pushups. I assume that the
mechanism is itself a control system, but at some level decisions have to be
made as to which perceptions must be controlled at a given moment and which
must be allowed to "float." My wife had a serious auto accident one morning
years ago when the "finding something in her purse" perceptual control
system became active while the "driving the car" perceptual control system
went on hold. Some higher-level system determines which of these goals will
be pursued at any given moment, and it seems to me that more is involved in
deactivating a system than just setting its reference level to zero. The
effect in some cases is as if the system ceases to exist--or as if the
PERCEPTUAL INPUT is set to the reference level.

I think you will find that "switching
mechanisms" are a hangover from the days of viewing organisms as cute, furry
computer systems. Even if there is data that actually seems to require a
model that switches higher level access to a lower level control system, the
switching system will almost certainly have to be a control system itself.

I don't think of MYSELF as a cute, furry computer system. Cute, maybe, but
not furry. But I think that we agree that the "switching mechanism" will
have to be a control system; I didn't mean to imply otherwise.

I can easily write a simulation that switches control systems in response to
changing input. I don't see any conceptual roadblock to the notion that we
living control systems can do what that simulation does.

Regards,

Bruce

[From Rick Marken (950302.0930)]

Dennis Delprato (950301?) --

NINTH ANNUAL BEHAVIOR ANALYSIS OF MICHIGAN CONFERENCE
(with Bruce Abbott on "What Does Perceptual Control Theory
Have to Offer the Experimental Analysis of Behavior?"

That sounds very interesting! What does PCT have to offer EAB, Bruce? So far,
it has apparently been an offer they _can_ refuse :wink:

Bruce Abbott (950301.2040 EST) --

This little game reminds me of the old game of "hot and cold." Move this
way and everyone says "you're getting hotter," move that way and they say
"no, cold, very cold!" The goal, of course, is to find whatever is hidden
and if the (negative) feedback is good, you inevitably will. One property
of "reinforcement" (and "punishment" as well).

Are you saying that "good" negative feedback is a "property" of
"reinforcement" and "punishment"?

For the magazine-trained rat, the click of the feeder is a signal that food
has arrived in the food cup; so control passes immediately to theposition-
sensing perceptual system with reference set to the cup's
coordinates.

Saying that "control passes immediately to the position-sensing perceptual
system" implies that the rat was not controlling sensed position before the
feeder click. There is a way to test this, of course. Just apply disturbances
to the rats position before and after the click and see whether the
disturbances are more effective before the click. I bet they are not.

I believe that a proper PCT-based model will account for the effects we
naive behaviorists refer to as reinforcement and punishment,

We already know it does. Bill's model of a rat shock experiment showed
that "punishment" is the value of a perceptual variable that differs
substantially from the reference state of that variable (or it is a
disturbance--like a shock-- that moves the perception from the reference
state). Bill's model of operant scheduling data shows that "reinforcement" is
the reference state of a controlled perceptual variable (or a disturbance--
like food-- that moves that variable to the reference state).

effectively explain so-called conditioned reinforcement and punishment,

Yes. I'm sure it can but I don't know of any working PCT models of this
phenomenon yet. It probably involves learning to control a variable (such as
the relationship between click and position relative to bar) in order to
control another variable (such as the amount of time spent pressing the bar
to get food).

and show why the term "self-reinforcement" is absurd.

Actually, it doesn't sound that absurd to me. If reinforcement is the
reference state of a controlled variable, then the state of that variable
(such as rate of food input) is being controlled in order to produce a result
(reference rate of food input) that the organism wants for _itself_. The term
"self-reinforcement" may be a bit vague, but I don't think it's absurd. It's
certainly not as absurd as the term "reinforcement" itself, which suggests
that there are events in the environment that can "control" (strengthen)
behavior.

But there are still a lot of gaps to fill, particularly how to deal with
what in operant terminology is called "stimulus control."

Now _there_ is a term that is absurd. It refers to an impossible entity -- a
"stimulus that controls". PCT deals with this kind of "stimulus control"
very easily; it says that it doesn't exist. What does exist (and what people
are probably describing when they use the term "stimulus control") is
"response to a disturbance of a controlled variable".

What mechanism selects which control system will have current access to the
lower-level systems required to produce the necessary behavioral outputs?

Why do you think that such a mechanism is necessary? If you run my
spreadsheet model you will see that the same set of lower level control
systems can be used _simultaneously_ as the means of controlling several
different higher level variables. I think you will find that "switching
mechanisms" are a hangover from the days of viewing organisms as cute, furry
computer systems. Even if there is data that actually seems to require a
model that switches higher level access to a lower level control system, the
switching system will almost certainly have to be a control system itself.

By the way, any news about the response of your colleagues to the
compensatory tracking data analysis?

Best

Rick

[From Rick Marken (950302.2140)]

Bruce Abbott (950302.1600 EST) --

Seriously, I'm planning to introduce PCT, list its similarities to EAB
(single-subject approach, etc.), point out at least one major difference
(behavior: control of perception not stimulus: control of behavior),
illustrate the basic control-system model, then apply the model in the
context of an asthma self-management program designed to teach
asthmatics how to control their asthma symptoms. I hope to show
that PCT provides a better model of the self-management process than
does one based on Skinnerian principles.

Sounds great! Break a leg!

The term "self-reinforcement" is usually applied in the context of
trying to establish control over some perception, such as one's weight.

Oh, I see. Yes, this kind of "self-reinforcement" is what I would call
controlling by setting up an internal conflict. You try to control a
variable (weight) by setting up a conflict beween two control systems in
yourself; the one that wants to eat and the one that wants to see a
movie. You do this by setting up a higher level control systems that
controls a contingency -- if <no food> , then <movie> else <no movie>.
This contingency control system is in conflict with other systems that
are setting goals for food and movie consumption. This approach to
controlling weight will work for a while, but the persistant error will
lead to reorganization that is seen as "giving in to temptation". So, this
kind of "self-reinforcement" is not really absurd; it's just painful
because you are putting youself in a position where you are purposely
losing control (of food intake and movie going) in the hopes that a
side-effect will be weight reduction.

Note how different the situation is if some outside agency determines
access to the movie.

It is somewhat different. When you control the perceived contingency
yourself you can "get around it" by changing your own reference for that
contingency. But you can't make someone else change their reference
for the contingency. When another person establishes the contingency
between food and movies you are still in conflict -- but now the conflict
is the result of the contingency created by an external agent. The solution
to such a conflict is to revise one's goals (for food and movies) and "play
along" with the contingency or (more likely) work around the contingency
(lie and cheat) or eliminate the person enforing it.

A pigeon is trained to peck at a response key (a plastic disk on the
wall) when the key is illuminated in red but to turn in a complete
circle when the key is illuminated in green. Either way the response,
when completed, pproduces brief access to a hopper full of grain. We
operant guys would say that the two responses, keypecking and
turning, are both under "stimulus control,"

Well, you guys will just have to get over that;-)

meaning that the probability of observing each response depends on
which stimulus (red or green) is currently present.

But that's not control. Why not call it what it is: a relationship between
variables?

What controlled variable is being disturbed?

My guess is that the controlled variable is the perceived relationship
between light color and action; the color of the key (red or green) is a
disturbance that is compensated by the actions (pecking, circling) of
the organism.

Why does behavior change when the "discriminative stimulus"
changes?

I know you know this but I'll play along. The behavior changes to
maintain the perception of the relationship "if red light then peck; if
green then turn in circle". If the red light came on and the bird did not
peck or if it turned in a circle it would not be perceiving the intended
relationship.

Does PCT really say that this corelation between observed behavior
and keycolor does not exist?

Now I know you're shinin' me on. Of course it says the correlation
exists becuase it exists. It says that the correlation (a relationship
perception) exists because it's under control.

A more productive approach would be to offer an explanation for it
from within the PCT framework.

I just did. Sure hope I passed the audition;-)

By asserting that stimulus control is "response to a disturbance to a
controlled variable," you are only stating the basic PCT assumption
that all behavior is error-driven.

No. I am stating that it looks like stimuli cause the responses of a control
system if you don't see the variable being controlled. Once you see that
the relationship between light color and response type is being controlled,
you see that the appearance of stimulius control is just that -- an
appearance (or what in my brasher days I would have called an illusion).

Best

Rick

[From Bruce Abbott (950303.2100 EST)]

Bill Powers (950302.1200 MST)

Thanks for the details on programming the e. coli-style reorganization
process; we had talked of this before but it was nice to see it laid out
concisely. In fact, I found the entire post extremely interesting, and the
one following it.

We need a superordinate system that can
turn off the outputs from one control system and connect another control
system that is already organized to use the same lower-order locomotion
systems to go get the food (from any starting position).

Yes, and I can visualize how that might happen, by inhibiting and
disinhibiting appropriate neural inputs. It's quite clear to me that we
construct new control systems all the time (isn't that what we mean when we
say that we've learned how to DO something?) and then just call them as
needed, like subroutines. When these systems become automatic enough we
call them "habits" or "skills." Of course, the "subroutine calls" are just
outputs of higher-level control systems.

An event would be defined as a set of such vectors existing between two
times, t1 and tm. I suppose there would be a way to represent such an
event as a scalar perceptual signal, but we can leave that problem for
the future (the only reason for doing so would be theoretical, anyway).
The question then becomes what picks out an event from the continuous
stream of variable values. We might guess that events are separated by
zero entries -- nothing happening. Or perhaps they are marked off by
particular perceptions occurring. You can see why I'm reluctant to get
very far into this -- about two steps of this sort of conjecture define
a year's worth of experimentation. And we haven't even asked how an
event would be _recognized_.

There may be innate "rules" to help identify which perceptual events should
be associated. Garcia's learned taste-aversion studies showed, for example,
that a rat that had been given "bright, noisy, tasty water" (sips of water
with an unfamiliar flavor that was delivered along with a flashing light and
beeping tone) and then made ill by injection with lithium chloride
subsequently avoided water of that flavor but not unflavored water delivered
with the flashes and beeps. But "bright, noisy, tasty water" followed by
footshock produced avoidance of "bright, noisy water" but not of the
flavored water. In other words, rats associated flavor with illness, light
and sound with footshock. Also, a traumatic, single event may become
associated with a wide range of perceptual inputs: not just a simple,
discrete stimulus that immediately preceded it but general visual,
olfactory, and auditory inputs surrounding the event, which help to identify
the entire context in which the event occurred. Then there are other
"rules" such as sequence (causes before effects) and temporal proximity to
the event which usually play a part. (This by no means exhausts the list.)

This is beginning to look to me like the question of how the
relationship level of control gets organized, just as the previous
comments look like asking how the event-level gets organized. This is
suggesting to me that there may be specific learning processes
associated with the hierarchical levels. Reorganization would come into
the picture only when no fixed method could create a control process. If
we inherit the basic machinery for developing these different levels of
control, it's not unreasonable to suppose that the method of learning
such levels is actually what is inherited. So we have some excuse to
look for possible methods.

This is one of those things I was trying to get across during the "e. coli
wars." Reorganization may be required when all else fails, but it would be
far more efficient to apply more systematic methods when the situation
permits. This is what my "learning" e. coli was doing, and I presume that
living control systems have evolved (and may learn) a set of strategies for
gaining control which may be followed initially.

That fits with the concept of perceiving correlations as one type of
relationship. A single instance means, and should mean, little. Only
when there is a consistent relationship over a period of time should we
believe that it is real. So the mechanism for creating relationship-
perceiving systems should only gradually create an automatic perception
of relationship when there are covariances.

This is what I found with e. coli: if the change in tumble probability
following a single tumble was too great, the system failed to adapt--it just
kept "changing its mind" about what to do and thus never developed an
effective, consistent strategy (i.e., tumble if nutrient decreasing, inhibit
tumble if nutrient increasing).

    However, other perceptual systems may have control--the smells, the
    cracks and crevices in the corner of the chamber, the view of
    things lying out of reach in the pan below the grid floor--and
    control may return to these systems after the food has been eaten.

Let's try to keep the language straight: perceptual systems can't
control anything.

Oops. I meant to say "perceptual CONTROL systems" there.

Getting the food may have temporarily reduced the error in the system
involved in the former activity, leaving other errors that are larger
and entail different actions to reduce them. But with enough
observations we ought to be able to get an idea of what the major
control systems are. What does a rat do during its busy day? Not too
many different things, I would think.

Yes, or some new perceptual input may have caused a higher-level control
system to activate a different lower-level control system by producing a
large enough disturbance. As to what rats do with their time, not much in
the sterile environment of an operant chamber (though perhaps surprisingly
more than you might expect), but in a"real life" setting it would take a
very long list indeed to include all the rat's control systems.

    But there are still a lot of gaps to fill, particularly how to deal
    with what in operant terminology is called "stimulus control."
    What mechanism selects which control system will have current
    access to the lower-level systems required to produce the necessary
    behavioral outputs? On what basis?

A lot of this would simply fall out of a sufficiently complex model, in
which many systems are controlling for many reference conditions at the
same time.

Yes, I can see that; this is, I think, what Rick Marken was proposing.

If you see the whole hierarchy as a collection of control systems all
trying to correct their own errors at the same time, you can get an
inkling of what the final picture will look like. All the interactions
will seek states in which overall error is minimized. As learning takes
place throughout the hierarchy, this minimum will gradually, over months
and years, get smaller and smaller until some irreducible amount of
error remains. If this minimum is too high, you get an anxious screwed-
up organism. If it's exceptionally low, you get a happy effective rat or
person.

Yep, that's what the world needs--happy, effective rats. D'ya ever wonder
what it would be like to study some nice, simple system like rocket
propulsion or nuclear synthesis? Man, those physicists have it made!

Bill Powers (950303.0945 MST)

I think we need to take these questions seriously. We have never tried
to model the situation where one perception appears to work as a signal
rather than as a controlled variable in itself. It's possible, of
course, that this interpretation of the role of the perception is
misleading, but we need to find the "correct" description and show that
it makes at least as much intuitive sense. The problem is similar to
that of showing that some "stimuli" should really be interpreted as
disturbances. Once you can see exactly what is disturbed, and how the
control action counteracts the disturbance and _appears_ to be caused by
the stimulus, the PCT interpretation becomes at least as believable as
the other one. We need to do this for the case of "discriminative
stimuli" or else admit that some stimuli serve as triggers for other
control processes.

I'm very glad to hear this, first because the concept of "discriminative
stimulus" ranks second only to "reinforcement" in importance within
traditional behaviorist theory and thus needs to be at least addressed by
PCT and second because it means you agree with me that it is an empirical
issue within PCT that requires serious attention and--ta da!--research. It
is also the subject of some very confused thinking (as I view it) within the
experimental analysis of behavior. This will require a bit of explanation.

The concept first arose in the context of providing a simple cue (such as a
light) to signal that a reinforcement contingency was now in effect: Light
on, lever-pressing activates feeder on, say, a VR-10 schedule; light off,
lever-pressing fails to activate feeder (extinction schedule). The rats
learned to lever-press when the light was on and to do something else when
the light was off. This observation required that a new term be invented
and defined:

Discriminative stimulus: A stimulus in the presence of which a response is
reinforced.

Along comes a new situation: green light = VR-10, red light = extinction.
The green light is the discriminative stimulus. What's the red light? So
we need a new term invented and defined:

S-delta: A stimulus in the presence of which a response is NOT reinforced.
(The "delta" should be given in superscript as the greek letter that looks
like a triangle.) In contrast to S-delta, we'll call the discriminative
stimulus "S-D" (S superscript D).

But in the first situation, wasn't S-delta the ABSENCE of S-D? Well, yes.
So the ABSENCE of a stimulus is a stimulus? Must be. Hmmmm. And shouldn't
BOTH S-D and S-delta be called discriminative stimuli? Looks like we need a
new definition of descriminative stimulus:

Discriminative stimulus: a stimulus that sets the occasion for a response.

That seems to work better: Green light = press the lever; red light = go do
something else.

More trouble: Green light = VR-25 shock delivery, red light = no shock;
FR-5 food delivery in either case. Responding is suppressed during green
(but still occurs), recovers during red. What's the green light? The red
light? Who's on first? Why didn't I take up something easier like nuclear
physics or fractal geometry?

Well, despite the confusion, everyone seems to know what everyone else MEANS
by these terms, even if they evade precise definition. These stimuli
indicate what contingencies (relationships) are "in effect" at any given
moment, and the well-trained rat's behavior changes instantly (if that is
what is required) when these stimuli change. The changes may be in the
goals being pursued (earning food, investigating the chamber) or in the
methods by which the goals are pursued (lever pressing, turning, both for
the same access to food).

There are two contributing lower-order perceptions under the organism's
control (peck, turn-in-circle) and two that are independently variable
(Red, Green). As I interpret the description, there are two additional
relationships imposed by the environment: red XOR green is true, meaning
that the light is either green or red but never both, and peck XOR turn-
in-circle is true, meaning it is physically impossible to do both at
once.

Correct.

The only way to find out what logical function (if any that we can
understand using Boolean algebra) describes the perceptual function is
to test various hypotheses -- the good old Test again. The actual
behavior involved, pecking or turning in a circle, is of little interest
in itself, however striking it may be to the experimenter. What we would
be investigating would be the logical function of perceptual variables
that is under control by the organism.

To do this it is also necessary to look carefully at the logic of the
experiment. What is the contingency for the case when both lights are
on, or both off, or when a light of a new color is shown? If the animal
turns in circles, pecking at the key once each time around, what is the
contingency? It's very easy to set up the logic of an experiment with a
particular set of relationships in mind, and forget that you have to
cover all combinations of true and false, not just the ones you first
thought of. Bill Leach will no doubt support this observation;
forgetting to cover all logical possibilities is a pitfall of electronic
logic design. Whichever condition you forgot to provide for is almost
certain to be the next one that occurs.

Easy to forget the other relationships? Like overlooking the fact that, if
light-on is the discriminative stimulus, that light-off must be something,
too? Now who would ever do that? (;->

Start with, say, a vertical line projected on a pigeon's response key as S-D
(reinforcement) and a horizontal line as S-delta (extinction), then train
the pigeon until it responds on the key only during S-D. What happens if
you now project a line on the key at a 45 degree angle? Answer: an
intermediate rate of keypecking. In fact, the rate will vary smoothly and
continuously as you vary the line angle smoothly and continuously, from zero
rate at horizontal to max rate at vertical and back to zero when you reach
horizontal again. The plot of this relationship is called a gradient of
generalization. There's a huge literature on this which includes some
surprises, but I think I've said enough for now. Do I sense another
research project coming on? (:->

Regards,

Bruce