[From Bruce Abbott (941011.1520 EST)]
In the year before B:CP was published (1972), Robert Rescorla and Alan Wagner
offered a model of classical conditioning that has come to be called,
logically enough, the Rescorla-Wagner model. It turned out to have
considerable power in that it correctly described a number of well-known
classical conditioning phenomena, including acquisition, extinction,
overshadowing, blocking, discrimination learning, and inhibitory conditioning,
and predicted a novel effect--overexpectation--that no one had thought to test
for. This prediction was subsequently confirmed.
The model is not without its flaws--it does not take the rather important
parameter of time into account, for example--but it did a remarkable job of
making sense of a number of well-established phenomena. Here is the model in
its simplest form, set up to account for the typical acquisition curve:
V the current associative "strength" of the connection between the
conditioned stimulus (CS) and unconditioned stimulus (US).
a the "salience" of the CS (its attention-gettingness), represented as a
number between 0 (no salience) and 1 (maximum salience).
B a learning-rate parameter proportional to the magnitude of the
unconditioned response to the US, with 0 representing no response (used
to model extinction trials) and 1 the maximum response.
dV the change in associative strength between CS and US following a
Pavlovian trial.
L the maximum conditioned response sustainable by the CS; the limit of
conditioned responding.
The model for simple conditioning on a single trial is:
dV = aB(L - V)
V = V + dV
Iterated across a series of trials, it produces a negative exponential curve
that rises to an asymptote at L. More salient CSs and more potent USs
increase the rate of acquisition. Applied to salivary conditioning, it says
that the amount of salivation to the CS should increase over trials to the
limit determined by L, with each subsequent pairing of CS and US producing
less of an increase than those before it.
An instructive question (for me anyway) is can this model be viewed as an
instance of feedback-regulated control? We seem to have the basic elements.
We might consider L as the set point of the salivary-control system regulating
the amount of salivation to the CS. V is the current amount of salivation
being produced by the CS; thus (L - V) functions as the comparator, with a and
B providing a slowing factor.
Let us now condition the maximum response to the CS (a tone) over a series of
trials, and do the same with a second CS (a light) in a separate series of
trials. What happens if we now present BOTH CSs together as a compound CS?
Rescorla and Wagner assumed that the associative value of a compound CS is
simply the sum of the values of the component simple CSs. If each component
CS is fully conditioned, their associative values are both close to L; combine
them and the value of the compound is close to 2L. The model predicts that
(1) initial presentation of the CSs as compound will produce a CS larger than
either stimulus is capable of producing by itself, and (2) that further trials
in which we CONTINUE TO PAIR the compound CS with the US (normally thought of
as a reinforcement procedure) will produce a REDUCTION in response strength
with a limit at L. These predictions were empirically confirmed.
A control-theory interpretation would simply note that the compound CS now
produces output above set point. The (temporary) disturbance produced by
combining fully conditioned simple CSs into a single compound produces a
negatively-signed error which gradually (over trials) weakens the associative
strength of the compound CS to bring it back to its reference state.
Now for the instructive part. Does this interpretation fail in its attempt to
apply PCT principles to Pavlovian conditioning? Explain.
Bruce (;-})