Pitch perception; scaley observer

[From Bill Powers (960513.1400 MDT)]

Martin Taylor 960513 10:20 --

     Take a harmonic complex, and add x Hz to each component, so that,
     say, a 100 Hz base complex restricted to the harmonics 500, 600,
     ... 1000 is replaced by a sound composed by adding together
     sinusoids of frequencies 512, 612, ... 1012. A lot of pitch
     perception work has been done on such sounds in the last couple of
     decades. This unnatural sound still has difference frequencies
     that are all harmonics of 100 Hz, and if the ear were using the
     nonlinear difference frequencies, that would be the pitch that is
     heard. But it isn't. This pitch that is heard is higher by an
     amount proportional but not equal to the added number of Hz (say
     the same as a 104 Hz harmonic complex, in this example--a number I
     pulled out of a hat rather than from experiment).

This result might be an encouragement to look further at the
harmonically-related phase-locked loops. In a series like 512, 612
...1012, the difference frequencies all are 100 Hz, but the "harmonics"
are not actually harmonically related. If the control signals from a
number of harmonically related synchronous detectors, or tunable
filters, were summed together, the higher harmonic contributions would
all indicate too high a frequency in comparison with the common 100-Hz
beat frequency, so the average result would be biased toward a higher
frequency, by an amount depending on the weighting given to the signal
from each harmonic detector.

I should mention that my so-called phase-locked loops were not actually
the standard design. They consisted of a tunable filter, with the input
and output of the filter multiplied together and smoothed to give the
output signal. In a tuned filter, the input and output are in phase only
at the resonant frequency. My design for a tunable filter (see nest
paragraph) provides two output signals 90 degrees out of phase. At
resonance, one output is in phase with the input, the other 90 degrees
of out of phase. When the 90-degree shifted output is multiplied by the
input to the filter, you end up with zero output from the multiplier
relative to the average. The output signal then goes negative (from some
average value) when the frequency rises, and positive when it falls.
This signal, amplified and integrated, is used to vary the tuning
parameter of the filter, and this is how the filter is made to track the
frequency of the input. The signal that varies the tuning parameter is
then used as the perception of frequency. Obviously you can get any kind
of nonlinear response you want, by putting the right function between
the error signal and the tuning parameter.

The tunable filter consists simply of two integrators in a closed loop,
one with a negative sign on its input. By adjusting the leakiness of the
integrators, you can select any Q -- sharpness -- for the resonant
filter. The gain associated with each integrator determines the
frequency; if both gains are varied together, the frequency is
proportional to the gain. If only one gain is changed, frequency is
proportional the square root of the adjustable gain.

This kind of filter is in principle realizable with two neurons and the
output multiplier with two more. Of course one would expect a real
neural circuit to use more than the minimum possible number of neurons.

Just sort of getting this design on the record, in case I never get back
to it and someone else wants to give it a try.

RE: detection of audio signals by electrodes

     Even though the _rate_ of firing doesn't change with the signal
     frequency, nevertheless, the _moment_ of most probable firing does.
     Up to a significantly high frequency (Peter says about 8 kHz), if a
     nerve fires, it will be at a moment phase-locked to a peak of the
     signal waveform. The thing is, that the higher the frequency, the
     less probable it is that any particular nerve will fire at any
     particular peak, so that over a few milliseconds averaging period,
     roughly the same number of firings will have occurred no matter
     what the frequency.

So: when an audio input is present, some of the firings bunch up near
the peaks of the audio waveform, and presumably are relatively spread
apart at the valleys. Thus while the mean rate of firing can remain the
same, there is a modulation of the mean rate at the audio frequency, so
a representation in terms of impulses per second would look like a mean
value plus an A.C. component.

If I understand what you're describing correctly, this is exactly how I
have been thinking of audio neural signals. The mean rate of firing
corresponds (with suitable nonlinearities) to intensity or rather
loudness, and the superimposed envelope of firing rates carries
frequency information -- in fact, carries the audio frequency signal as
neurally represented, like this( 3 peaks shown):

/ / / / / //// / / / / ///// / / / / / ///// / / /

Now the mean rate of firing could remain exactly the same while the
frequency component changes, and vice versa. Actually, in non-ideal
neurons we would expect some interaction between mean rate of firing and
the superimposed AC signal.

At the intensity-control level (for example the reflex that tightens and
loosens the muscle connected to the eardrum, analogous to the iris), the
above signal would be integrated over enough time to erase the frequency
modulation, and the average signal would represent the intensity to be
controlled (as in an automatic gain control). The same signal, sent on
to a frequency-discriminating level of perception, would be treated as
in my tunable filters, where the tuning neurons would have an adjustable
integrating response of a speed appropriate to the frequency modulation
of the spike rate. The integration factor depends on the rate at which
incoming impulses bump the post-synaptic potential upward, and the
leakiness depends on how fast the potential decays between input
impulses. The output frequency of relaxation oscillations from the
neural integrator is set from moment to moment by the post-synaptic
potential.

According to Rosenblatt's Principle, the signal diagramed above would
NOT constitute a perception of frequency. It is analogous to the RF
signal in my example of yesterday, in which the audio frequency is
implicit in the signal but has not yet been turned into an explicit
representation of the audio signal. What is needed is a signal with a
spike rate that is proportional to the frequency of dense patches in the
above diagram. My tunable-filters approach is one (rather elaborate) way
of doing this; the signal that adjusts the tuning is a spike train whose
frequency is proportional to the FM modulation frequency, suitable for
use as a perceptual signal whose spike rate stands for the frequency of
modulation. There are simpler ways to do this, analogous to a rectifier
followed by a leaky integrator.

The basic signal diagrammed above is a _pressure_ signal, at the level
of intensities where pressure is represented by spike rate. As this
pressure varies, the spike rate varies. To turn this signal into a
representation of frequency, we need a function that can receive this
pressure signal and emit a signal whose spike rate is proportional to
the pressure-modulation frequency. Or, if the method of representation
is of the mapping variety, we need a function that will represent
different modulation frequencies as activation of sensory neurons in
different places in a map. Whatever the actual mode of representation of
frequency, it will occur farther up in the audio chain than the
intensity signals, and will contain signals which explicitly encode
frequencies (rather, pitches).

This makes me ask, how far up the audio processing chain did those who
looked for spike rates proportional to frequency look? Of course if the
representation is a place in a map, the question is irrelevant.

···

-------------------------
     "Subjective experience" to me has the qualities of an external
     analyst who is observing the actions of such a hierarchy.

I think we are simply going to differ on this question.

If a person has to point to "green" on a spectrum,
control is easiest if there is a single signal for green that can be
specified.

     What would "more green" and "less green" signify? Would "less
     green" mean "less saturated same spectral peak", "spectral peak
     nearer _this_ value" (a two-sided optimization problem), "less
     intensity in this spectral region", or what?

This depends on the kind of representation you're talking about. If
signal magnitude (frequency) represents color on a continuum (my second
proposal which you didn't mention), then "green" is simply one magnitude
of a signal that represents a range of colors. If we have place mapping,
then we have a number of color perceivers activated to varying degrees,
and something that receives their outputs would have to construct a
signal representing place in a spectrum.

     But it _is_ the analyst, if you are referring to "subjective
     experience" as the criterion. ... It's when one labels or
     verbalizes the nature of a perception that it seems probable that a
     single-valued scalar is entailed.

Once again, we differ. This sort of comment makes me wonder if you have
any experiences of the world which are not accompanied by names or
verbalizations. Actually, I presume you do, but they don't seem to play
any part in your discussions.

     And that is precisely what the "generalized flip-flop" connection
     (usually) achieves at what I call the "category interface" in the
     hierarchy.

Show me. Otherwise, no comment.

     The label "a perception" is presumably used because there is some
     perception to which it refers--and that perception defines a CEV,
     and therefore the "real world" must be so constructed that the CEV
     has some reality.

Why must a CEV "have some reality?" Why can't at least some CEVs amount
to order arbitrarily imposed on the inputs?

      By using the words "become a perception" you seem to intend to
     imply that something in the real world exists, and is properly
     labelled "a perception."

Not my intent. My intent was to assert that one perception goes with one
scalar signal. Until a complex input has been reduced to a set of simple
scalar signals, no perception takes place. I'm not saying that is right,
but that's what I intended my words to convey.

     (I have a wording difficulty here, in that I accept each scalar
     perceptual signal as the output of a single PIF in a network, and
     would use the word "a perception" to refer to its value, yet this
     "perception" is not the perceptual unity to which you refer, such
     as "green" or "house," or "good- natured").

I make a distinction between a perception and a symbol we use to point
to a perception. When I say "green" I usually mean the perception that
is indicated when I use the word "green." The perception itself is not a
label. It is a signal; it is precisely that unity that we experience as
a perception. It is the very unity that makes me think of one-
dimensional signals. A basic quality of perception can change only in
its magnitude.

     The hierarchy works equally well provided that the signals at the
     different levels are sufficiently different to allow the necessary
     control actions to control the perceptual signals. The Observer to
     whom you refer is what I called the "external analyst" who is free
     to take any complex in the world observed and to extract from it a
     single degree of freedom as a scalar value.

The first sentence is self-evident, but it doesn't answer the question
of _how_ these signals are different, and exactly what differences are
required to allow controlling the perceptions. Each brain ends up with a
specific set of perceptions, not a generalized field of possibilities.
In order for anything to happen, the brain must commit itself to a
specific organization. Also, I propose that the brain _does_ extract
single degrees of freedom to control them.

If you're going to spell Observer with a capital O, then the Observer is
not any kind of "analyst." The observer with a small o can indeed
extract single degrees of freedom and control them individually or
independently and in parallel. This observer is simply the set of
higher-order control systems.

I wonder what the alternative to extraction and control of single
degrees of freedom is. Is there really any alternative? I know that, for
example, you might decide to represent the operation of some set of
control systems in terms of vector control: _O_ = F(_R_ - _P_), where
the underlined variables are vectors or even matrices, and F is a matrix
operation. But if you ever had to implement this matrix control process
in a physical system (for example, in a computer), you would find that
you have to expand all the vector and matrix notations into the
elementary processes that carry them out.

     To see why it doesn't matter to a PIF in the hierarchy whether
     lower-level perceptions are labellable, just think of the simplest
     form of PIF, a weighted summation. If the inputs are labelled
     scalars, the exact same sum can come as from any linear rotation in
     the space of the inputs, with appropriately changed weights. The
     set of outputs of PIFs at level N is totally unaffected by whether
     the inputs from level N-1 are discretely labelled or are
     distributed across the whole set.

I agree that labels don't matter at the lower levels which are not
concerned with controlling for labels. But you speak as if the input
weightings are continually and freely variable, which I very much doubt.
Arriving at a workable set of weightings takes a long time, and once
such a set is found it may not change for the rest of your life. For a
given person, a given level of perceptions consists of approximately
orthogonal ways of perceiving the world. This way of perceiving may be
unique to that person and be shared by no other. And I doubt that there
are many changes in it once it has reached a sufficient state of overall
internal consistency -- not without rather drastic upheavals.

     The only difference of opinion that I detect is whether an external
     observer could always take one signal value and identify it with
     the value of something the observer sees as a unitary "thing" in
     the outer world.

This is not the difference between us. I claim that control systems have
simple comparators that deal with scalar variables, not complex
variables. And I claim that whatever we experience, the experience
consists of scalar neural signals. None of this has anything to do with
the external correlates of neural signals. I am claiming, rightly or
wrongly, that this is how the nervous system becomes connected. If you
want to say "Well, this is only a proposal and it could be wrong," I
will agree immediately. But this is the proposal I find most reasonable
right now.

     There's no analysis in a linear rotation of a basis space, and no
     value in having more levels than a single one that extracts (as
     scalar signals) combinations that are useful to control.

You're going to have to explain that one. Are you saying that there is
only one level in the hierarchy? Are you saying that mentally rotating
an image does not change any perceptions of it (such as its
orientation?).

     My comment was to the effect that one should not impose a
     designer's convenience _a priori_ on the investigation of a system
     that has evolved for no convenience of an external analyst or
     designer.

I guess I have to ask you this. Do not designers use the same brain
functions that the rest of us use? Is the designer's convenience any
different from the brain's convenience? You seem to speak of external
analysts or designers as if they operated according to some principles
different from those that apply to the brain they are thinking about. Is
not the external analyst simply demonstrating some of the behaviors of
which the brain is capable?
-----------------------------------------------------------------------
Kent McClelland (960513, direct) --

Kent, we're delighted that you're coming to the meeting, and looking
forward to hearing about the current status of your work "on creating
a PCT sociology." Your registration has been received. We are going to
try to double your money in the Lottery. OK?
-----------------------------------------------------------------------
Best to all,

Bill P.

[From Peter Cariani (960514.1200 EDT)]

[From Bill Powers (960513.1400 MDT)]
Martin Taylor 960513 10:20 --
     Take a harmonic complex, and add x Hz to each component, so that,
     say, a 100 Hz base complex restricted to the harmonics 500, 600,
     ... 1000 is replaced by a sound composed by adding together
     sinusoids of frequencies 512, 612, ... 1012. A lot of pitch
     perception work has been done on such sounds in the last couple of
     decades. This unnatural sound still has difference frequencies
     that are all harmonics of 100 Hz, and if the ear were using the
     nonlinear difference frequencies, that would be the pitch that is
     heard. But it isn't. This pitch that is heard is higher by an
     amount proportional but not equal to the added number of Hz (say
     the same as a 104 Hz harmonic complex, in this example--a number I
     pulled out of a hat rather than from experiment).
This result might be an encouragement to look further at the
harmonically-related phase-locked loops. In a series like 512, 612
...1012, the difference frequencies all are 100 Hz, but the "harmonics"
are not actually harmonically related. If the control signals from a
number of harmonically related synchronous detectors, or tunable
filters, were summed together, the higher harmonic contributions would
all indicate too high a frequency in comparison with the common 100-Hz
beat frequency, so the average result would be biased toward a higher
frequency, by an amount depending on the weighting given to the signal
from each harmonic detector.

These "pitch shift" experiments with inharmonic complex tones led to the
realization that whatever the auditory system is doing, it must do a
fine structure analysis of the waveform and/or the power spectrum rather
than an analysis of waveform envelope and/or frequency spacings. The pitch
shift simply "drops out" of an autocorrelation analysis (as de Boer showed
in his 1956 thesis), but it can also be accounted for in spectral terms,
the major models being Goldstein's model, which is based on best matches to
"spectral templates" that are harmonic series, and Terhardt's model, which
is based on a "subharmonic sieve". There are now also neural net models
based on these kinds of strategies. As I understand your proposal, the
organism or device would be adaptively generating harmonically-structured
signals, in effect doing adaptive filtering,
in order to find the best fit to a harmonic series. I like the
proposal because it fills a gap in the existing strategies, that mostly
assume pre-existing pattern-recognition mechanisms (either learned or evolved).
It appears to me to be almost an "analysis-by-synthesis" strategy which
is adaptively controlled. I've thought about these general kinds of
strategies in neural terms (sets of assemblies that generate particular
time patterns that are then cross-correlated with incoming time patterns).
The major problem that one would encounter in a neural implementation would
be realizing precise harmonic relations (if implemented using a "place-pattern"
generators) or generating precise time patterns (if operating using temporal
pattern generators).

I should mention that my so-called phase-locked loops were not actually
the standard design. They consisted of a tunable filter, with the input
and output of the filter multiplied together and smoothed to give the
output signal. In a tuned filter, the input and output are in phase only
at the resonant frequency. My design for a tunable filter (see nest
paragraph) provides two output signals 90 degrees out of phase. At
resonance, one output is in phase with the input, the other 90 degrees
of out of phase. When the 90-degree shifted output is multiplied by the
input to the filter, you end up with zero output from the multiplier
relative to the average. The output signal then goes negative (from some
average value) when the frequency rises, and positive when it falls.
This signal, amplified and integrated, is used to vary the tuning
parameter of the filter, and this is how the filter is made to track the
frequency of the input. The signal that varies the tuning parameter is
then used as the perception of frequency. Obviously you can get any kind
of nonlinear response you want, by putting the right function between
the error signal and the tuning parameter.

The tunable filter consists simply of two integrators in a closed loop,
one with a negative sign on its input. By adjusting the leakiness of the
integrators, you can select any Q -- sharpness -- for the resonant
filter. The gain associated with each integrator determines the
frequency; if both gains are varied together, the frequency is
proportional to the gain. If only one gain is changed, frequency is
proportional the square root of the adjustable gain.

This kind of filter is in principle realizable with two neurons and the
output multiplier with two more. Of course one would expect a real
neural circuit to use more than the minimum possible number of neurons.

There are correspondences between the operations that you are describing
and the auto- and cross-correlation procedures I outlined the other day, and
there are some psychophysical models that combine time and "place"
information using "matched filters" in the auditory pathway (Goldstein &
Srulovicz's model for frequency selectivity). It's worth noting that
neurally one can do all sorts of operations with phase delays using
simple delay lines and coincidence detectors (and there are many sensory
systems that use this strategy). There is a very good paper on these
kinds of neuro-computations in bats by Simmons et al in the book
"Auditory Computation", 1996, Hawkins et al, eds., Springer-Verlag.
The one problem with the tunable-filter approach proposed above
(in terms of human pitch perception) is that human perception of the
pitch of complex tones is highly invariant with respect to the phase
of the components (except for a few very specific and subtle effects when
higher-frequency harmonics are used). You could use a bank of these filters to
zero in on the frequency components in the signal (after band-pass filtering
has been done), but then you need to analyze that patterns of output
signals in each channel in a manner that is not dependent upon their
relative phases. (This might be accomplished using a rich set of delays and taking
sets of "best delays", which is equivalent to the "delay compensation"
that was in the cross-correlation procedure that I outlined, and then using
another layer of such filters). The advantage of using autocorrelation
(or neurally of using all-order interspike interval distributions)
is that phase is eliminated in the representation itself (Simmons et al
also grapple with this problem in the above article.)

     Even though the _rate_ of firing doesn't change with the signal
     frequency, nevertheless, the _moment_ of most probable firing does.
     Up to a significantly high frequency (Peter says about 8 kHz), if a
     nerve fires, it will be at a moment phase-locked to a peak of the
     signal waveform. The thing is, that the higher the frequency, the
     less probable it is that any particular nerve will fire at any
     particular peak, so that over a few milliseconds averaging period,
     roughly the same number of firings will have occurred no matter
     what the frequency.

I should clarify about the upper limits of phase-locking. Based on mammalian
auditory nerve data and human evoked auditory potentials, phase-locking is
thought to be very strong up to about 1-2 kHz, declining rapidly as
frequency increases, so that the upper limit of phase-locking in humans
is thought to be about 4-5 kHz. Some animals that use auditory timing
cues very effectively have upper limits of phase locking that are higher
(I think the barn owl is highest, up to 9 kHz). These figures are for
continuous pure tones; for other kinds of temporally-structured stimuli (noise,
click onsets, etc.) there is usable temporal information about the onsets
of transients and low-frequency periodicities in frequency channels above these limits.

So: when an audio input is present, some of the firings bunch up near
the peaks of the audio waveform, and presumably are relatively spread
apart at the valleys. Thus while the mean rate of firing can remain the
same, there is a modulation of the mean rate at the audio frequency, so
a representation in terms of impulses per second would look like a mean
value plus an A.C. component.

This is mostly correct, although talking in terms of a "mean rate" in this
context (500 usec windows for a 1 kHz tone) is a bit misleading. It's the
AC component that carries the information, either in its magnitude ("temporal-
place" information, AC + DC) or in its periodicity ("interval" information).
I've never liked the term "instantaneous rate" because it has always seemed to
me to be oxymoronic: "rates" always need to be computed over some (contiguous)
time interval. "Periodicity" on the other hand is a function of the relative
times-of-arrival of sets of events (a joint property of time points,
not contiguous in time). I never use the word "frequency coding" because it
most often is assumed to mean "rate-coding", as in the number of events in
a given time window, rather than "periodicity coding", the distribution of
intervals or the power spectrum of the signal.

If I understand what you're describing correctly, this is exactly how I
have been thinking of audio neural signals. The mean rate of firing
corresponds (with suitable nonlinearities) to intensity or rather
loudness, and the superimposed envelope of firing rates carries
frequency information -- in fact, carries the audio frequency signal as
neurally represented, like this( 3 peaks shown):

/ / / / / //// / / / / ///// / / / / / ///// / / /

Now the mean rate of firing could remain exactly the same while the
frequency component changes, and vice versa. Actually, in non-ideal
neurons we would expect some interaction between mean rate of firing and
the superimposed AC signal.

The two components, AC and DC, largely covary in the auditory nerve, so it
is difficult to separate out the representational consequences of the
two kinds of information.

At the intensity-control level (for example the reflex that tightens and
loosens the muscle connected to the eardrum, analogous to the iris), the
above signal would be integrated over enough time to erase the frequency
modulation, and the average signal would represent the intensity to be
controlled (as in an automatic gain control). The same signal, sent on
to a frequency-discriminating level of perception, would be treated as
in my tunable filters, where the tuning neurons would have an adjustable
integrating response of a speed appropriate to the frequency modulation
of the spike rate. The integration factor depends on the rate at which
incoming impulses bump the post-synaptic potential upward, and the
leakiness depends on how fast the potential decays between input
impulses. The output frequency of relaxation oscillations from the
neural integrator is set from moment to moment by the post-synaptic
potential.

The problem here is that the efferent systems of the auditory system,
in the middle ear and in the cochlea, are much slower (tens to hundreds
of milliseconds and longer) than the intensity fluctuations you want to
control. It's also still a matter of debate how much compensatory gain
these systems have, but they appear to be fairly weak.......In an artificial
system, these limitations need not apply......

This makes me ask, how far up the audio processing chain did those who
looked for spike rates proportional to frequency look? Of course if the
representation is a place in a map, the question is irrelevant.

There are "locker neurons" in auditory thalamus that follow stimulus
periodicities up to 1 kHz, and (a very few) similar kinds of units have
been seen at the level of the primary cortex in awake animals, but again
this looks much more like a "periodicity code" than a "rate-frequency"
code. Those people who have looked for cortical "pitch detectors" would
have seen "rate-frequency" units, since they generally are paying attention
to neural discharge rates, although this negative finding needs to be
strongly tempered by the very few numbers of people who have looked
and the relatively few places that they have looked.

(Powers reply to Martin Taylor):

     My comment was to the effect that one should not impose a
     designer's convenience _a priori_ on the investigation of a system
     that has evolved for no convenience of an external analyst or
     designer.

I guess I have to ask you this. Do not designers use the same brain
functions that the rest of us use? Is the designer's convenience any
different from the brain's convenience? You seem to speak of external
analysts or designers as if they operated according to some principles
different from those that apply to the brain they are thinking about. Is
not the external analyst simply demonstrating some of the behaviors of
which the brain is capable?

I think the point here is that we tend to conceive of the brain (which
we don't understand) as working something like our most advanced
technological artifacts (in this case, computers, neural nets, or control systems,
which we think we understand), but there may be principles involved
in its functioning that we haven't discovered yet.
We need to be careful not to project our own current technological artefacts
onto the brain so strongly that they obscure other processes that might be
going on. At the same time, we can get ideas about new kinds of artefacts
if we examine the brain with sufficiently open minds.

In response to a post by another CSGer:
Although control theory need not use multiplexed, multidimensional signals and
can operate on scalars, it is conceivable that these kinds of signals could
permit more powerful and flexible ways of organizing control networks (so that
this discussion of "scalar vs multidimensional" signals, and rate vs. time
codes <is> relevant to PCT in general and to PCT explanations of brain function
in particular-- obviously, "crap" is in the eyes of the beholder).

Peter Cariani

[Martin Taylor 960514 22:05]

Bill Powers (960513.1400 MDT)

This is very short--mostly to let you know that I will be incommunicado
for about 10 days. But there's some kind of major misunderstanding
(not uncommon) between us, since I cannot reconcile the following:

Me:

     The only difference of opinion that I detect is whether an external
     observer could always take one signal value and identify it with
     the value of something the observer sees as a unitary "thing" in
     the outer world.
Powers:
This is not the difference between us. I claim that control systems have
simple comparators that deal with scalar variables, not complex
variables. And I claim that whatever we experience, the experience
consists of scalar neural signals. None of this has anything to do with
the external correlates of neural signals. I am claiming, rightly or
wrongly, that this is how the nervous system becomes connected. If you
want to say "Well, this is only a proposal and it could be wrong," I
will agree immediately. But this is the proposal I find most reasonable
right now.

···

--------------

Everything I said in the message on which Bill P was commenting was
predicated on these same beliefs. These are _not_ differences of opinion
between us.

Martin