Rescorla-Wagner and PCT

[From Bruce Abbott (941011.1520 EST)]

In the year before B:CP was published (1972), Robert Rescorla and Alan Wagner
offered a model of classical conditioning that has come to be called,
logically enough, the Rescorla-Wagner model. It turned out to have
considerable power in that it correctly described a number of well-known
classical conditioning phenomena, including acquisition, extinction,
overshadowing, blocking, discrimination learning, and inhibitory conditioning,
and predicted a novel effect--overexpectation--that no one had thought to test
for. This prediction was subsequently confirmed.

The model is not without its flaws--it does not take the rather important
parameter of time into account, for example--but it did a remarkable job of
making sense of a number of well-established phenomena. Here is the model in
its simplest form, set up to account for the typical acquisition curve:

V the current associative "strength" of the connection between the
     conditioned stimulus (CS) and unconditioned stimulus (US).

a the "salience" of the CS (its attention-gettingness), represented as a
     number between 0 (no salience) and 1 (maximum salience).

B a learning-rate parameter proportional to the magnitude of the
     unconditioned response to the US, with 0 representing no response (used
     to model extinction trials) and 1 the maximum response.

dV the change in associative strength between CS and US following a
     Pavlovian trial.

L the maximum conditioned response sustainable by the CS; the limit of
     conditioned responding.

The model for simple conditioning on a single trial is:

                    dV = aB(L - V)
                     V = V + dV

Iterated across a series of trials, it produces a negative exponential curve
that rises to an asymptote at L. More salient CSs and more potent USs
increase the rate of acquisition. Applied to salivary conditioning, it says
that the amount of salivation to the CS should increase over trials to the
limit determined by L, with each subsequent pairing of CS and US producing
less of an increase than those before it.

An instructive question (for me anyway) is can this model be viewed as an
instance of feedback-regulated control? We seem to have the basic elements.
We might consider L as the set point of the salivary-control system regulating
the amount of salivation to the CS. V is the current amount of salivation
being produced by the CS; thus (L - V) functions as the comparator, with a and
B providing a slowing factor.

Let us now condition the maximum response to the CS (a tone) over a series of
trials, and do the same with a second CS (a light) in a separate series of
trials. What happens if we now present BOTH CSs together as a compound CS?

Rescorla and Wagner assumed that the associative value of a compound CS is
simply the sum of the values of the component simple CSs. If each component
CS is fully conditioned, their associative values are both close to L; combine
them and the value of the compound is close to 2L. The model predicts that
(1) initial presentation of the CSs as compound will produce a CS larger than
either stimulus is capable of producing by itself, and (2) that further trials
in which we CONTINUE TO PAIR the compound CS with the US (normally thought of
as a reinforcement procedure) will produce a REDUCTION in response strength
with a limit at L. These predictions were empirically confirmed.

A control-theory interpretation would simply note that the compound CS now
produces output above set point. The (temporary) disturbance produced by
combining fully conditioned simple CSs into a single compound produces a
negatively-signed error which gradually (over trials) weakens the associative
strength of the compound CS to bring it back to its reference state.

Now for the instructive part. Does this interpretation fail in its attempt to
apply PCT principles to Pavlovian conditioning? Explain.

Bruce (;-})

Tom Bourbon [941011.1658]

[From Bruce Abbott (941011.1520 EST)]

In the year before B:CP was published (1972), Robert Rescorla and Alan Wagner
offered a model of classical conditioning that has come to be called,
logically enough, the Rescorla-Wagner model. . . . Here is the
model in its simplest form, set up to account for the typical acquisition
curve:

Am I right that you imply the model takes different forms to explain
different phenomena? If so, it is un-PCT-like.

V the current associative "strength" of the connection between the
    conditioned stimulus (CS) and unconditioned stimulus (US).

Can you tell us the units in which "associative 'strength'" is measured?
How does the experimenter determine the present value?

a the "salience" of the CS (its attention-gettingness), represented as a
    number between 0 (no salience) and 1 (maximum salience).

Units? How does the experimenter determine the present value?

B a learning-rate parameter proportional to the magnitude of the
    unconditioned response to the US, with 0 representing no response (used
    to model extinction trials) and 1 the maximum response.

Units = units of saliva? leading to a unitless ratio?

dV the change in associative strength between CS and US following a
    Pavlovian trial.

Units? How does the experimenter determine the present value?

L the maximum conditioned response sustainable by the CS; the limit of
    conditioned responding.

How does the experimenter determine the value at the limit?

The model for simple conditioning on a single trial is:

                   dV = aB(L - V)
                    V = V + dV

Iterated across a series of trials, it produces a negative exponential curve
that rises to an asymptote at L. More salient CSs and more potent USs
increase the rate of acquisition.

How does an experimenter determine the "salience" or the "potency" or both?
What happens to behavior, moment-by-moment, during a "trial?" Does the
model make any attempt to model momentary behavior, rather than the results
of behavior integrated over rather large "chunks" of time -- trials?
. . .

An instructive question (for me anyway) is can this model be viewed as an
instance of feedback-regulated control? We seem to have the basic elements.
We might consider L as the set point of the salivary-control system regulating
the amount of salivation to the CS. V is the current amount of salivation
being produced by the CS; thus (L - V) functions as the comparator, with a and
B providing a slowing factor.

You say that L might be a set point, but for which perceptual signal? As
you describe it here, L seems to specify the magnitude of a response, not of
a perception. Or do you mean that the animal controls the sensed amount of
salivation? Do I mis-read you here? Might it be the case that the level
of salivation is not controlled, but that the animal controls something
like sensations of moistness, or sensations that pertain to the consistency
of food? If the animal controls one of those alternative perceptions,
rather than the amount of salivation, then the "objectively measured" amount
of salivation at any moment is an uncontrolled variable by means of which
the animal controls perceptions of something else.

It seems that we can translate the phrase, "V = the current associative
"strength" of the connection between the conditioned stimulus (CS) and
unconditioned stimulus (US)," into the simpler and clearer phrase, "V = the
amount of salivation when the CS is present." The definition of dV would
also become simpler.

Using the simpler definitions of V and dV, I have a few questions about the
conditioning experiment, as you summarized it.

Let us now condition the maximum response to the CS (a tone) over a series of
trials, and do the same with a second CS (a light) in a separate series of
trials. What happens if we now present BOTH CSs together as a compound CS?

Rescorla and Wagner assumed that the associative value of a compound CS is
simply the sum of the values of the component simple CSs.

Would this mean, in other words, "the amount of salivation when there is a
compound CS is the sum of the amounts when either simple CS is presented
alone?"

If each component
CS is fully conditioned, their associative values are both close to L; combine
them and the value of the compound is close to 2L. The model predicts that
(1) initial presentation of the CSs as compound will produce a CS larger than
either stimulus is capable of producing by itself, and (2) that further trials
in which we CONTINUE TO PAIR the compound CS with the US (normally thought of
as a reinforcement procedure) will produce a REDUCTION in response strength
with a limit at L. These predictions were empirically confirmed.

A control-theory interpretation would simply note that the compound CS now
produces output above set point.

The (temporary) disturbance produced by
combining fully conditioned simple CSs into a single compound produces a
negatively-signed error which gradually (over trials) weakens the associative
strength of the compound CS to bring it back to its reference state.

I'm not sure that is a PCT interpretation. For one thing, I don't think
many of us would say "the compound CS produces output." The reasons for my
pickiness over the wording should be obvious.

Might it be the case, for example, that each of two CSs has come to be
associated its own perceptual function, either of which can act through the
same salivary apparatus to affect the real (as yet undetermined) controlled
variable? (For the moment, I will not discuss the association between, for
example, visual perceptual functions and perceptual functions for the
controlled variable(s).) Either CS alone will produce error signals that
activate the salivary apparatus, bringing the actual controlled variable
under control. When the two CSs occur together for the first time as a
compound CS, their independent perceptual functions produce error signals
that sum and produce a high level of output from the salivary apparatus,
which drives the actual controlled variable to a state that produces error
relative to its reference level. The rest would be simple PCT, with the
amount of salivation declining as an uncontrolled variable by means of
which the organism controls the as-yet-unidentified controlled variable.

Now for the instructive part. Does this interpretation fail in its attempt to
apply PCT principles to Pavlovian conditioning?

I believe it does.

Explain.

I tried.

I believe an effective PCT model would produce actions that achieve
perceptual control of something other than salivation, with "trials"
representing relatively large "chunks" comprising many moments, and with
salivation representing an action by which the organism controls some other
variable(s).

Later,

Tom