Reinforcement, PCT style

[From Bruce Abbott (951205.1320 EST)]

Bill Powers (951204.1520 MST) --

    Bruce Abbott (951204.1550 EST)

    When the effect of a disturbance is such that the contingent
    delivery of the so-called reinforcer would not reduce this error,
    it would not act as a reinforcer; in fact if its effect were now to
    _increase_ the error, it would actually appear to _suppress_ the
    behavior that produced it.

I don't think you've quite got the picture.

I think I do get the picture, but I need to communicate it better. What you
describe is quite right; if the event we are terming the reinforcer reduces
the error, the output necessarily declines. If there is a steady
disturbance to the system, the output will stablilize at an equilibrium
value that will not completely cancel the error; but will just balance the
error-increasing effect of the disturbance. That much I understand.

The reinforcer is the effect of the control system's action on the
controlled variable. In operant experiments it is usually delivered in
quanta of a definite size (e.g., 45 mg pellet) immediately following the
action that produces it. [The size given in the example is only
proportional to the quantity of whatever the controlled variable is that the
pellet supplies.]

I am maintaining that if the reinforcer did not tend to reduce error, it
would not serve as a reinforcer. Within the control system, a constant
disturbance would tend to increase the error; the system will reach an
equilibrium in which the reduction in error supplied by the reinforcer will
just balance the increase in error supplied by the disturbance, so one will
not _observe_ any further reduction in error once this state is reached;
nevertheless, the reinforcer continues to supply an error-reducing effect on
the controlled variable each time it occurs as a result of control-system
action. If it ceased to provide this service, it would cease to function as
a reinforcer.

Now, if a disturbance to the system could push it into an opposite error
(controlled variable above reference rather than below it), the same
"reinforcer" would now contribute to an exacerbation of the error rather
than to its reduction. A consequence of action that did so would serve, not
as a reinforcer, but as a punisher. Not only would the system stop
producing the action that produces the "reinforcer," it would actively
oppose any other system tending to produce that action (even as a
side-effect!), so long as that action continued to produce the same effect
on the controlled variable as before. The two systems would be in conflict.

This effect will not be seen with food unless the action delivers the food
directly into the rat's gut. If the action (pressing the lever) merely
causes a pellet to be delivered into the chamber, this by itself does not
affect the error in the rat's system, assuming that the controlled variable
is something like the rat's nutrient state or fullness of the gut. The
effectiveness of the reinforcer in maintaining lever-pressing in the normal
situation comes about because it keeps a whole chain going: lever-press
-----> pellet ------> approach to pellet, picking up pellet, consuming
pellet. When the pellet becomes aversive (increases error when consumed),
the link between lever-pressing and change in controlled variable is broken,
so suppression of lever-press action will not occur.

The situation you describe in which the reinforcer reduces error and
therefore the system output would be described in reinforcement theory in
terms of drive and satiation. A disturbance (food deprivation) arouses
error ("hunger drive") which generates output proportional to error
(intensity of behavior proportional to strength of drive). The action
produces an effect (consequence) on the controlled variable that tends to
reduce error (reinforcer); reduction in error (reduced drive) reduces the
output (lowered intensity of behavior) until the error is reduced to zero
(if there is no continuing disturbance); the approach to this state is
termed "satiation."

To determine whether a given consequence of behavior serves under present
conditions as a reinforcer, one must assess whether the intensity of a given
action is higher when the action produces the putative reinforcer than when
the action does not produce it. This is very different from comparing the
intensity of action before and after a control system has adjusted to
counter a disturbance, which is the comparison you were making. This is why
I said that we seemed to be on different wavelengths.

Regards,

Bruce