[From Bill Powers (960123.0600 MST)]
Bruce Abbott (960122.2115 EST) --
This is where all the stuff about some representation of the
responses being added into short-term memory comes into play. As
each "response" occurs (whether "target" or "other"), its
representation is added to STM; and the representations of
previously experienced responses "decays" to some fraction of their
previous values. The current representation of the "target"
response is the sum of all these decaying representations. When a
target response is finally followed by an incentive, the current
representations of all responses (both target and other) are
"associated" with the incentive delivery according to the current
strengths in memory. The strength of this association is reflected
in the "coupling" between the incentive and the various responses,
both target and nontarget. Killeen says that the incentive
delivery, by filling STM with representations of the consummatory
responses, tends to effectively remove the target and other
responses from STM so that their representations are essentially
zero after the incentive has been dealt with.
So what does the representation of a response look like in short term
memory? A response consists of exerting a downward force and then
releasing it. What is it about this act that is stored? And what is it
that decays? Is it the amount of remembered force? The duration? The
direction? That is, as the memory decays, does the amount of remembered
force or its duration or its direction decay? Is the sum of all
remembered stored forces greater than any one of them? Or is it
something else about the responses that is stored? And what good does it
do to store remembered measures of a response?
I have nothing against invoking memory where it seems necessary -- for
example, in explaining how an animal goes to its cache of food when it's
hungry, or finds a bone where it was buried last year, and so forth. But
I would much rather suppose that perceptions, not actions, are
remembered; what good does it do to remember an action when the
environment has changed and a different action is required to get the
same perception as before?
Killeen offers an elaborate hypothetical mechanism involving response
memory, including interactions among memories and special effects of
reinforcements on erasing memories -- a very complicated scheme, the
only result of which is an exponental term in an equation. What he's
doing sounds an awful lot like the sorts of things personality
psychologists do that gets them criticized by EABers.
M is multiplied by rho, the proportion of responses that are target
responses, to give the steady-state value of zeta, the coupling
coefficient. By entering target response values as 1.0 and
nontarget response values as 0.0, one can arrive at the value of
zeta directly in the iterative version of the equation for M, which
is just our old buddy the leaky integrator: M = beta*y + (1 -
beta)*M, where M is initialized at zero after each incentive
delivery and y is either 1.0 or 0.0 depending on whether the
current "response" is or is not a target response.
The point of this whole fairy tale about response memory is just to
arrive at an equation with the form necessary to make the curve bend
down where it should. Wouldn't it have been simpler to say "We need a
term that rises to an asymptote from zero as N increases from 0" and
then just write it? Many different forms might have been used other than
an exponential; other forms might even fit the data better. Staddon
proposed ellipses for the whole curve; I seem to remember that someone
else proposed parabolas. The best form would be a table representing the
actual observed function, deferring guesses about its explanation.
I like your assessment:
The real problem for Killeen's scheme is a familiar one: defining
"responses." Killeen tells us how to compute the "representation"
of the response but does not really tell us what a "response" is.
He seems to view behavior as a continuous stream that may or may
not include the instrumental act ("target" response) but does not
tell us how to parse that stream. And like most behaviorists he
seems to take the response as some repeatable (and repeated)
movement or sequence of movements, apparently unaware of the
difficulties variable environmental disturbances pose for such an
analysis.
So far so good (though I dispute that Killeen told us how to compute the
representation of a response), but the following is the critical error:
So long as one is dealing with a repeating _consequence_ of varying
behavioral acts, one can treat all acts having this consequence as
instances of the "same" target response, perhaps even narrowing the
definition a bit by observing the minimal repeating acts required
to produce the consequence.
The difficulty with this "class of responses" idea is that the class
must include both a response and its opposite, and any degree of
response between those limits. The motor response to the sight of a
piece of food might have to be turning right or turning left, pushing or
pulling, lifting or pressing down. Or it might have to include doing
nothing, if a disturbance provides the required result without the need
for any action, especially when any action would prevent the result from
occurring properly (as in driving a car around a curve in a crosswind).
It should be a sign of trouble when one has to say that the response an
animal learns is "anything it has to do to create the right effect."
This puts a smudge in the middle of the logical argument, leaving
unanswered the question of how the animal just happens, on each
occasion, to pick the right magnitude of the one variant of the behavior
that will produce the right result, where all other magnitudes and
variants would have produced the wrong result. How would you make a
model in which one box is labelled "right response generator?" What
would you put in that box?
The discriminative stimulus is an attempt to handle this problem. But
what do you do when there is no possible discriminative stimulus, yet
the right variant of the response is still produced every time? This is
the case when disturbances act directly on the consequence without at
the same time providing in advance any separate sensory cues. The
success or failure of a model turns on its ability to explain this case,
the most common one. The only answer I have ever seen to this problem
(outside PCT) is the claim that some sort of discriminative stimulus
MUST have been present, even though we didn't observe it -- some kind of
"subtle cue." In other words, the only answer is to deny that the
problem exists.
Thus, once the pigeon is standing within a comfortable striking
distance of the key, "responses" can be defined as the cycle of
forward head-thrust, leading to the striking of the key with
sufficient force to close the switch, and the return movement that
carries the head back to its initial position. But this analysis
fails to explain why the pigeon moves to the key in the first place
and would use a different set of muscles to strike the key from a
different angle if required to do so by the placement of some
partial obstruction between the pigeon and the key.
Precisely. You relieve me, sir. This is the crux of the difference in
approaches.
The basic notion behind Killeen's scheme is that organisms will
tend to repeat whatever they were doing at the time an incentive
suddenly appears, with the "whatever" being encoded as a series of
act memories which are rapidly fading as new acts enter STM. The
notion that they will "do again" that which they "did before" seems
a reasonable one, but "that-which-is-done" needs to be redefined in
terms of reference signals of control systems at appropriate
levels. The pigeon is pecking the key, not because some incentive
has forced it to emit a forward-thrust of the head, but because it
wants to repeat the act of striking the key with its beak.
This is still too limited. The main difficulty is the stereotype of
behavior which underlies all experiments of this kind; it is built into
the experiments. The "act" that the animal must produce is always
conceived as a unitary event, something that simply happens at a point
in time. The idea that it might occur over a continuum of magnitudes and
directions simply doesn't come up, because the very apparatus is set up
so it registers only an event, the closure of a contact. Under such a
conception of behavior, the ONLY dimensions in which behavior can change
are place and rate. This leads to a very limited idea of what behavior
is: it is the repetition of acts in certain places at certain rates.
I have been rather disappointed by the pace at which our plans for
experimentation have been going (not your doing, I understand). What I
had hoped for was to go through some basic replications of simple op-
cond experiments, and then start introducing more kinds of behavioral
variables. What I envision, for example, is a rat holding a wire lever
in its paw and moving it up and down to maintain a pointer on a mark,
with food delivery being contingent on its maintaining the match for a
few seconds. Or perhaps the rat could be in a running wheel, with the
speed of the wheel determining the pointer position. Or perhaps the rat
could press on a strain gage to position the pointer. Maybe the rat
could learn to turn a crank, to make a moving pointer match the speed of
another moving pointer, or maybe it could adjust the proportions of one
rectangle to match the changing proportions of another. A rat might
learn to balance itself on an unstable platform. The possibilities, once
you get beyond rate and place of repetitive acts, are endless, and we
would stand to learn much more about rat organization than otherwise.
What kind of variables can rats control, and how well? Could we discern
levels of rat organization?
An important theoretical question to resolve is how the delivery of
the incentive leads to the establishment of this reference.
Aw, Bruce! How about the theoretical question of what establishes the
reference for the incentive itself, turning an ordinary sensory
experience into a controlled variable?
···
-----------------------------------------------------------------------
Remi Cote 220196.1615 --
In B:CP (p.83) there is a drawing of a first-order control system.
The comparator is the spinal motor neurone cell-body. The
reference signal is incoming from the top from another axon. It is
represented with a + sign (my guess is that it is a kind of
activator). The Golgi tendon receptor give an input to the same
comparator (represented by a - (minus)). If there is more (+) than
(-) in the motor neurone cell body, there is a firing of this
neurone.
I guess this is the basic. Do I understand it well? That is the
question
If you're like most people, you skipped Chapter 3, which is about neural
currents and neurones as computing elements. The main concept you need
is that neural signals are measured as frequencies of firing, not as
individual impulses. When you see a + sign at a synapse, this means that
the incoming signal causes an increase in the outgoing signal: as the
input frequency increases, the output frequency increases. There is no
direct correspondence between incoming impulses (in this kind of neuron)
and outgoing impulses. As the incoming impulses arrive, the
concentration of certain ions inside the neuron rises with each incoming
impulse and slowly decays, with the average concentration depending on
the rate at which the impulses are received. The average ion
concentration determines how fast the neuron recovers after each firing
and fires again, so ion concentration is turned into an output frequency
of firing. The output frequency can be quite different from the input
frequency -- either greater or less, but proportional to it.
A synapse with a minus sign results in _decreasing_ the ion
concentration inside the neuron, so the faster the negative or
inhibitory impulses arrive, the more slowly the neuron will fire
(assuming it is also receiving positive or excitatory impulses at the
same time). If the neuron is receiving both positive and negative
inputs, its net rate of firing will be approximately proportional to the
difference in the positive and negative input frequencies. If the
negative input just equals the positive input in terms of net effect on
internal ion concentration, the output frequency will just be zero. As
Bill Leach pointed out, this works only when the positive input
frequency is greater than the negative input frequency: when inhibition
predominates, the neuron simply doesn't fire.
If you look on page 28 you'll see a neural subtractor. As Bill Leach
said, this is the simplest possible comparator. We usually say that the
reference signal has the positive sign, but depending on the external
connections in the control loop the comparator could also work with the
signs reversed. In Chapter 9, p. 117, you will see that in the "central
nuclei of cerebellum" this interchange of signs seems to exist in this
part of the brain, where the incoming signals from above are uniformly
inhibitory or negative in their effects, while the sensory feedback
signals are positive or excitatory -- just the opposite of the case for
the spinal tendon reflex systems.
I don't know if my understanding of the cerebellar connections is still
considered right.
----------------------------------------------------------------------
Best to all,
Bill P.