grandmother cells etc.

[From Bill Powers (960514.0600 MDT)]

     ... the concepts of low- vs. high-level perceptions bear some
     discussion. This gets into a whole nest of issues about perception.
     The way most machine perception people and neuroscientists think
     about it is that there are various "feature detectors" in the
     ascending pathway, so that the signals get simpler, but what they
     'represent' gets more complex. Although this idea of the
     "pontifical neuron" has been successfully satirized by Jerry
     Lettvin's example of the "grandmother cell", the notion of a
     feedforward hierarchy of processes is still tacitly held by almost
     everybody.

The "grandmother cell" which is specialized to recognize a grandmother
is still pretty low in the hierarchy in my scheme of things. Anyway,
that particular mocking satire, which has been unthinkingly repeated for
at least 30 years and has become one of the code-terms of cybernetics,
is silly on the face of it, just as silly as the claim that there are
cells in the brain that recognize orientations of lines. One cell can't
do anything very interesting. If we measure a neural signal that
correlates with the presence of something in the environment, that can't
be a sign that the cell generating it is itself is doing the
recognizing; we're just looking at the output of a computing process
that preceded this output. The signal is the result of the preceding
computations, even though the _state_ of the perception could be
indicated by the frequency of a single neural signal. I see nothing
unreasonable about the proposition that the perception we call "my
grandmother" boils down to the presence of a particular neural signal
(redundancy of pathways aside). The whole problem of perceptual
computation comes down to how that signal is generated.

It may be easier to understand my concept of hierarchical perception by
looking at an artificial perception: relative humidity. This measurement
is generated out of two temperature measurements, a dry-bulb and a wet-
bulb temperature. In an electronic humidity meter built on this
principle, each thermometer generates a signal that stands for a
temperature. The two signals are inputs to a computation that could work
in any of a number of ways. There could be a two-dimensional stored
table of values with the input signals addressing one entry in the
table, or there could be a symbol-handling computer evaluating an
analytical function with two arguments, or there could be an analog
computer that does the computation using the physical properties of
circuit components. However the computation is done, the result is a
single output that ranges from zero to 100 as indicated by the output
meter. Of course this signal has the right meaning only if the right
computations are done to combine the two readings; any other computation
would also generate an output signal that is a function of the two
temperatures, but only one function would produce a signal that varies
linearly with relative humidity.

In this hierarchical system, there are three perceptual signals: two at
the first level and one at the second. The perceptual signal at the
second level is the relative humidity signal; the signals at the first
level are temperature signals. If higher levels of circuitry were added
to this system, all three of these signals could in principle be
available to the higher systems for use in further computations, or
simply for monitoring. That is, a perception of relative humidity does
not preclude also perceiving air temperature in the normal way, through
the dry-bulb reading, or perceiving something like wind-chill
temperature through the wet-bulb reading. The same input information
could be used in several different ways by further layers of circuitry.

The problem with the cybernetic idea of the grandmother cell is in the
assumption that the recognition of "your grandmother" is the _only_
level of perception involved. If you go down a level, you might find
signals representing a host of attributes of your grandmother, such as
the color of her hair, the tone of her voice, the smell of her living-
room, the pet phrases she uses, the expression on her face, her height,
and so forth. Each one of these attributes could conceivably also be an
attribute of another person, but when they all exist at once, the result
can be a perception of only one person, your grandmother. And that
unique perception can be represented by a single signal which is present
when all the required attribute-signals are present at the same time
with the proper magnitudes. Of course the cell that finally emits this
single signal was responsible only for generating the weighted sum of
all the attribute signals; other computing networks at lower levels were
responsible for generating the attribute signals. Each attribute signal
is likely to be the output of even lower-order computing processes
receiving even more detailed perceptual signals.

     There are other general classes of possibilities, though. It is
     possible that every station has some access to an "analog"
     representation of the entire signal (e.g. in the form of temporal
     spike patterns present in the single neuron discharges of a
     population of cells), and that particular high-order invariances
     are extracted when needed by particular stations (like pitch or
     phonetic identity, or timbre, etc.) or when new aspects of the
     signal become relevant (when "perceptual learning" is needed).

This view is not completely different from my hierarchical model,
because in my hierarchical model higher systems can, in principle,
receive copies of the perceptual signals in any lower system. The main
difference is that in the arrangement you suggest, we never get down to
one signal representing one attribute EXCEPT BY INDIRECT IMPLICATION.
When you speak of "extracting" an invariance, just what do you mean? How
is a particular invariance represented once it is extracted? In your
example concerning pitch perception, you proposed a model to do the
extraction, and the final outcome was a normalized sum of a set of
harmonics -- a single scalar number. In order to explain how that number
can lead to further perceptual processes, or to behaviors like
controlling that scalar number, isn't it simplest to assume that the
number must be neurally represented?

My hierarchical model focuses on the process of extracting invariances,
rather than on the attributes of the signals from which they are
extracted. A perceptual function in my model is an invariance-extractor,
and the result of the extraction is an actual neural signal.

Notice that in the relative humidity detector, the percentage of the
maximum carrying capacity of air for water is implicit in the two
temperature readings. However, if we consider these two readings to be a
complex lower-level perception, it is clear that even though relative
humidity is somehow contained in this perception, it is not contained in
a usable form. That is, there is no way to convert these two temperature
readings into an action that will raise or lower the moisture being
added to the air, as a way of controlling relative humidity. Even though
the two signals may be varying as relative humidity varies, there is no
way to separate the humidity variations from the temperature variations.
This separation has to be done by a computation that generates a new
signal which is relatively insensitive to air temperature per se and
sensitive primarily to the proper function of the two temperatures. Once
that signal exists and varies with the relative humidity, control of
relative humidity becomes simple.

The same holds for control of pitch. Yes, a pitch-perception is an
imperfect representation of the fundamental frequency of a complex
waveform, but that simply results in control of an imperfect
representation (the system doing the controlling doesn't know it's
imperfect). The point is that there must BE a representation in order to
accomplish control of pitch in the simplest way.

In your proposal, a higher-level system would have access to the raw
input signal, and would extract invariances as appropriate to the higher
level. As an overall picture, I don't find this unreasonable. However,
as you state it this implies a great deal of duplication of functions,
because multiple higher-level systems which make use of similar but not
identical invariances have to do the entire process of extraction from
scratch. I think it is much more parsimonious to assume that there are
levels of invariance-extraction, providing multiple signals which can
then be used as inputs to higher systems without that level of
extraction having to be done again by any higher system. If a new basic
invariance is needed, it can be learned at the lower level, and then it
becomes available for many new kinds of higher-order processes. One of
the main advantages of the hierarchical arrangement is that it greatly
reduces the duplication of computing functions.

By the way, when you say "invariance" I hope you don't mean something
that doesn't vary. The size of a balloon is an invariance, but it can
obviously be altered. It is an invariance only in the sense that it
retains a specific identity: size. Its state, however, can vary.

···

------------------------------------
     We probably need to do some calibration of what the term "rate-
     place" code means. You were discussing the output of a pitch
     detector in which the rate of firing was montonically related to
     pitch frequency, (say, 50 spikes for 100 Hz, 100 spikes for 1000
     Hz, up to 200 spikes for 10 kHz). I don't know of any units like
     that anywhere.

I agree that we need to refine some meanings. The way you word this,
however, I think I detect a reference to cells in which rate of firing
does correspond to physical frequency over some range, although not over
the entire ranges you mention. This brings to mind the case of the
stretch receptors in muscles, of which essentially the same is true. In
any one fiber, a stretch signal will start rising in frequency at some
degree of stretch, and saturate at some higher amount of stretch. No one
stretch signal covers the entire range. However, when we consider all
the stretch signals that are associated with the same major muscle, and
add up the frequencies of firing, we find that not only is the group
rate of firing monotonic, but it is quite linear with muscle length as
well. I forget the reference for this, but I definitely read this in the
literature.

When you speak of a "unit" I assume you are speaking of signals in a
single fiber. Isn't it possible that a rate signal could exist in a
_group_ of fibers such that the group rate could be monotonic even
through the individual rates are not? In the spinal control systems, we
have many parallel systems, all sensing essentially the same physical
variable and contributing to the same physical output effect, so that
the control loop is really made of many partly-redundant loops. When I
think of "a perceptual signal," I am really thinking of a signal that
exists across a bundle of fibers, all of which carry the same kind of
information from one locality to another, but in many parallel paths. So
a signal that has a range of say 0 to 1000 spikes per second in a single
path could actually have a range of 0 to 100,000 spikes per second
reaching the same destination via 100 redundant pathways. Is this
physiologically realistic? If so, what would the real numbers be?

     There are, of course, many units in all auditory stations that are
     "frequency-tuned" in the sense that, for a given sound pressure
     level, they fire maximally at a particular frequency (so this is a
     non-monotonic rate-frequency curve). A "rate-place" code assumes
     that some central processor analyzes the rates in each "best
     frequency" channel (like in a spectrum analyzer that one has with
     some stereos or equalizers), and deduces properties of the stimulus
     on this basis.

When you say "frequency-tuned" I think of a tuning peak, with a sharp
reduction in response both above and below the center frequency. Is that
what you mean, or are you also including the case of saturation, where
the response rises over some range and then remains at maximum as the
input frequency continues to rise? If the latter, then we could be
speaking of a single fiber in a group which actually produces a
monotonic group firing rate over the whole range.

Has anyone tried to apply the perceptron approach to a set of frequency-
sorted signals? It seems to me that this ought to work. Given a set of
signals laid out in a lineal array, we have what amounts to a line-image
of a sound spectrum, quite analogous to an optical image. The problem of
pattern-recognition would then be the same, or even simpler since fewer
dimensions are involved. I find that in looking at an oscillograph-type
display of a vocal signal, I can easily distinguish some phonemes from
others just from the bumps in the pattern. Some are hard to distinguish,
but they might yield to a some way of filtering the signal before
display.

     The envelope approach doesn't explain human pitch perception
     (Schouten and de Boer in the 40's, 50's and 60's effectively
     falsified these mechanisms using inharmonic AM tones)

_Which_ envelope approach? It seems to me that there would be a zillion
ways to model a system using the envelope (rate-of-firing)
representation of neural impulses. I'm sure that Schouten et al
falsified ONE envelope approach, but certainly there would be others
they didn't test. The nature of the model depends entirely on the kinds
of computations you assume -- saying they have falsified "the" envelope
approach sound like a serious overstatement to me.

     I'm down on Dualism because of the Cartesian notion of separate
     "substances".

Well, if that's your only objection there's no problem. The only
"substance" I am talking about is neural tissue. For all I know, the
Observer is made of neural tissue.

     Here, also, the population-interval distributions could be the
     explicit signal. Relatively many 10 msec intervals in a particular
     population might well be the "representation for" a particular
     pitch.

Gee, this is hard to get across. I don't think you understand how weird
this statement sounds to an electronics engineer. If you don't somehow
extract and average those 10-msec intervals in some device, all you have
is a mess of different intervals. You need an INDICATION that there is a
large number of 10 msec intervals, a SIGNAL that is present when they
are present. How do YOU know that a tracing of a signal contains a lot
of 10-msec intervals? You have to look over the tracings, perhaps
counting the number of intervals of various sizes, or just squinting at
the whole thing and trying to get a sense of the average spacing. You
have to feed the raw data into some sort of interval-length detector.
Suppose you got interested in the number of 10-msec intervals that are
followed by a 20-msec interval. Would that suddenly make that
combination "salient?" No, you would have to DETECT the existence of
this pattern, with a pattern-detector tuned to respond to that pattern
and no other. It is only the response of the pattern-detector that tells
you there is any pattern there. In fact, the pattern you see depends
entirely on the pattern-detectors you apply to the raw data. There are
no patterns that are "really there."
--------------------------
     Problems of consciousness aside (which are really metaphysical
     problems) ...

No. Consciousness is an observational problem. The metaphysical problem
is in _explaining_ consciousness. I don't try to do that.
--------------------------
     ... when one finds very strong correspondences between patterns of
     neural activity and patterns of perceptual judgments (observed from
     the outside, by psychophysicists), and these cannot be explained in
     any other way, one is led to believe that these patterns of
     activity have something to do with the perception (and may even
     constitute the "neural code" underlying the perception).

Who says there are any patterns of neural activity, or any patterns of
perceptual judgements? An observer of those patterns, that's who.
Somebody is applying various pattern-detectors to observations and
seeing if there is any response. If there is a response, the observer
says that a pattern has been detected. If the observer forgets his own
role in detecting the pattern, of course, the observer might say that
the pattern is "just there," not realizing that it wouldn't be there if
he weren't equipped to detect it.

     I don't necessarily think that being an "observer" means being
     "consciously aware" -- maybe my ideas will change, I don't know.

I don't think that being an "observer" means being consciously aware,
either. Observing is simply processing neural input signals to produce a
neural output signal. All my experience tells me that a great deal of
this sort of observation goes on without my conscious awareness of it.
But my experience also shows me that there are many observations which
can take place either with or without conscious awareness, which is one
of the bits of evidence I interpret as meaning that awareness and
perception/observation are not the same thing. What is different about
breathing when I am aware of it, and when I am not? It obviously works
quite well without my conscious interference, but equally obviously I
CAN become aware of it, and I can even interfere with it rather
drastically, as in holding my breath, for brief times.

I use the term Observer, with a capital O, to refer to whatever it is
that can be selectively aware of neural signals that are already in
existence in the brain. I don't know what that something is.

     I'm just proposing accounts of observers and conscious awarenesses
     that are different from the "point-source receiver" model.

Don't take that "point-source" too literally. Basically, what I mean is
that ideas about spatial relations and spatial extension are activities
in the brain OF WHICH WE CAN BECOME AWARE. So you can think about where
awareness is or how big it is, but when you become aware of those
thoughts they are just thoughts, not whatever it is that is aware of
their existence. "Point-receiver" or "point-source" are just thoughts.

     Imagine a human hierarchy where "lower-level" functionaries get the
     primary signals, compile a report and send the report AND the
     primary signals upstairs for further analysis and decisionmaking.

That's what I did; that's how the hierarchy works.

     If the middle-level executives have no other concerns and the
     report looks consistent with what they expect, they go with the
     report, but if there is competing input or confusion, they may also
     do more data analysis or send out the primary signal + all of the
     accumulated reports to other consultants for evaluation. They might
     just sit tight and wait for more reports to come in...

The trouble with this way of thinking about it is that all these
executives have all the functions of a whole brain. You've gone beyond
the little man in the head problem and created a little crowd of men in
the head problem. Now you have to have a model for each of these little
men, explaining how things can look consistent, what confusion is, how
the inputs are converted to outputs, how each little man knows where the
others are and what they are doing and what the resulting effect on the
external world means and ... uh-uh. I don't buy it.

In the hierarchy I have proposed there is no little man in the head,
except the whole brain. The various levels accomplish the various
functions needed by the whole brain, each level specializing in one kind
of perception and control that uses lower levels and is used by higher
levels. There is no analogy with the way people interact with each
other, because no one level is a whole person. No level understands
anything about higher or lower levels. Only the whole brain is a whole
person.

     One of the problems I have with scalar codes is that they result in
     "combinatorial explosion" when one tries to deal with all of the
     combinations of explicit scalar percepts that are possible.

The way I deal with them generates no combinatorial explosion. The brain
does not handle all the combinations of explicit scalar percepts that
are possible. It handles only those it handles, only those that matter
to it. And when there are many combinations that are equivalent, one of
them is used, not all of them. Furthermore, large parts of my
hierarchical model deal with analog variables, to which combinatorics
does not apply. Your problem with combinatorial explosions is a hint
that you're using some inappropriate assumptions.

     Yes, one can do it, certainly, but is this the way it's done in the
     brain? We don't know, one way or the other, but at this point we
     need as many alternative hypotheses as we can get.

I agree that we have to consider many alternatives, but I think it
wisest to put them through some sort of plausibility test first.
-----------------------------------------------------------------------
Best,

Bill P.

[From Bruce Gregory (960414.1710 EDT)]

(Bill Powers 960514.0600 MDT)

Since Rick is on vacation, somebody will have to fill in for him.

Wow!

Thanks,

Bruce G.