Random echoes on pitch perception

[From Bill Powers (960511.1800 MDT)]

Peter Cariani, May 10, 1996 --

The above reference should have accompanied the initial citations in my
960510.2010.

In that same post from Peter, I find this:

     For spectrally-complex stimuli that give rise to low pitches, one
     with a fundamental frequency of 200 Hz will not yield more or fewer
     spikes than one at 220 Hz. The numbers of spikes produced are a
     function of the level, not the fundamental frequency!

That's fascinating. It suggests that the place where these discharges
are being measured is tapping an intensity (or loudness) perception
rather than a pitch perception. Where in the auditory perceptual chain
was this finding made? Or perhaps I should ask, how it is determined
that the signal being measured is itself the pitch perception, rather
than being a signal that is on the way to a pitch-perceiving system? Or
maybe even -- what sort of neural coding is assumed for perception of
loudness?

     Maybe you can explain to me how a 160 Hz pure tone, harmonics 6-20
     of a 160 Hz fundamental, AM noise with Fm=160 Hz, and a click train
     with a 160 Hz fundamental can all be matched within 1% to each
     other, using a neurally-plausible rate-based processing.

Do I take it that "harmonics 6-20" means that the fundamental and the
first five harmonics are missing, so that the lowest actual frequency
presented is 960 Hz? Got to think about that one....

Actually, if this input were passed through a nonlinear amplifier, there
would be frequency components in the output at the difference (as well
as sum) frequencies, and since there are 14 coherent contributors to the
160-Hz difference frequency, the output could actually contain a rather
strong 160-Hz component. This assumes that the harmonics are phase-
locked. The mechanical connections in the ear are somewhat nonlinear, so
the 160-Hz component could actually be present in the mechanical
vibrations reaching the cochlea. And of course neural responses are
nonlinear, too.

I don't know what Fm = 160 Hz means. Is this the lower frequency cutoff
of the noise signal, the modal frequency, or what? If my guess about
where the 160-Hz tone comes from in the second case is reasonable, all
but this noise-input case would seem to boil down to detecting the
lowest harmonic, the fundamental. That's not so tough. Well, not
impossible-sounding.

In fact, I did some experiments with phase-locked loops receiving
prolonged vowel sounds, precisely to try to see what sort of system
would be needed to detect the lowest frequency in a voice signal uttered
at different pitches. You can't define a formant in a general way unless
you can relate it to the fundamental, which is always changing in normal
speech.

A phase-locked loop tends to lock in not only on the fundamental, but on
higher harmonics (the first two or three have the strongest "aliasing"
effect). I found that if the frequency control system has a bias toward
zero frequency, the variable-frequency oscillator used as the reference
oscillator would start at a sub-audio frequency, which would then rise
quickly until the first lock-in occurred (in a tenth of a second or so).
After lock-in, the system would track very large changes in the
fundamental frequency without losing lock. This turned out to be a very
reliable way of identifying the fundamental, as long as the fundamental
was really there. And the signal driving the variable-frequency
oscillator can be used as a perceptual signal indicating the frequency.

My idea about the harmonically-related phase-locked loops came out of
these experiments, because sometimes the fundamental would be obscured
by noise, and various harmonics would get blurred out, too. So I figured
that if we had a phase-locked loop for each harmonic -- not as hard to
do as it sounds -- then the control of frequency could be based on a
sort of average error over all the harmonics, and one or two dropping
out (even the fundamental) wouldn't matter. I haven't actually tried
this one -- as usual, other things came up.

But I don't suppose you're interested in these ideas, since they are all
based on a frequency model of neural signals. A different train of
thought (although the same train of impulses!).

···

-----------------------------------------------------------------------
Best,

Bill P.

[From Peter Cariani (960513.2000)]

[From Bill Powers (960511.1800 MDT)]
In that same post from Peter, I find this:
     For spectrally-complex stimuli that give rise to low pitches, one
     with a fundamental frequency of 200 Hz will not yield more or fewer
     spikes than one at 220 Hz. The numbers of spikes produced are a
     function of the level, not the fundamental frequency!
That's fascinating. It suggests that the place where these discharges
are being measured is tapping an intensity (or loudness) perception
rather than a pitch perception.

You can think of it this way, as rate encoding an intensity in a particular
frequency channel, although there are even problems with this simple notion
at sound levels near threshold and at moderate to high sound levels
(re: Weber's Law)

Where in the auditory perceptual chain was this finding made?

I see this at the level of the auditory nerve and cochlear nucleus and
the rest of the literature shows similar kinds of behaviors at more
central auditory stations. As you go up the pathway, however, there is
more inhibition, more complicated spectrotemporal fields, more non-
monotonic rate-level functions, longer neural recovery times, so that
the simplified notion of a monotonic, nicely behaved frequency map
becomes rather strained by the time one reaches the primary auditory
cortices. People have looked for "pitch detectors" using many different
kinds of stimuli (pure and complex tones), usually looking at firing
rate, but nobody has found a map in which avg. firing rate covaries with
fundamental frequency over the entire range of fundamentals (say 100-
1000 Hz) or pure tone pitches (say 100-10000 Hz).

Or perhaps I should ask, how it is determined
that the signal being measured is itself the pitch perception, rather
than being a signal that is on the way to a pitch-perceiving system?

I think I explained how we think about this in terms of a putative neural
code or representation and how well (or whether) it always covaries with
the pattern of pitch judgments (that are externally observed
by psychophysicists).

Or maybe even -- what sort of neural coding is assumed for perception of
loudness?

Most models of loudness "count spikes", but as I say there are problems
with these, particularly at high levels. I think it's obvious that the
time structure comes into the loudness equation at some point, but this
is a bit involved.

     Maybe you can explain to me how a 160 Hz pure tone, harmonics 6-20
     of a 160 Hz fundamental, AM noise with Fm=160 Hz, and a click train
     with a 160 Hz fundamental can all be matched within 1% to each
     other, using a neurally-plausible rate-based processing.

Do I take it that "harmonics 6-20" means that the fundamental and the
first five harmonics are missing, so that the lowest actual frequency
presented is 960 Hz? Got to think about that one....

This is the case of the "missing fundamental".

Actually, if this input were passed through a nonlinear amplifier, there
would be frequency components in the output at the difference (as well
as sum) frequencies, and since there are 14 coherent contributors to the
160-Hz difference frequency, the output could actually contain a rather
strong 160-Hz component. This assumes that the harmonics are phase-
locked. The mechanical connections in the ear are somewhat nonlinear, so
the 160-Hz component could actually be present in the mechanical
vibrations reaching the cochlea. And of course neural responses are
nonlinear, too.

Nonlinearities can produce distortion products, but the missing fundamentals
are heard at levels where distortion products are very weak (inaudible).

I don't know what Fm = 160 Hz means. Is this the lower frequency cutoff
of the noise signal, the modal frequency, or what? If my guess about
where the 160-Hz tone comes from in the second case is reasonable, all
but this noise-input case would seem to boil down to detecting the
lowest harmonic, the fundamental. That's not so tough. Well, not
impossible-sounding.

Fm=160 Hz is the amplitude modulation frequency of the noise. The noise
envelope waxes and wanes every 6.25 msec and we hear a pitch at 160 Hz.
It's weak, but strong enough to carry a melody.

In fact, I did some experiments with phase-locked loops receiving
prolonged vowel sounds, precisely to try to see what sort of system
would be needed to detect the lowest frequency in a voice signal uttered
at different pitches. You can't define a formant in a general way unless
you can relate it to the fundamental, which is always changing in normal
speech.

A phase-locked loop tends to lock in not only on the fundamental, but on
higher harmonics (the first two or three have the strongest "aliasing"
effect). I found that if the frequency control system has a bias toward
zero frequency, the variable-frequency oscillator used as the reference
oscillator would start at a sub-audio frequency, which would then rise
quickly until the first lock-in occurred (in a tenth of a second or so).
After lock-in, the system would track very large changes in the
fundamental frequency without losing lock. This turned out to be a very
reliable way of identifying the fundamental, as long as the fundamental
was really there. And the signal driving the variable-frequency
oscillator can be used as a perceptual signal indicating the frequency.

My idea about the harmonically-related phase-locked loops came out of
these experiments, because sometimes the fundamental would be obscured
by noise, and various harmonics would get blurred out, too. So I figured
that if we had a phase-locked loop for each harmonic -- not as hard to
do as it sounds -- then the control of frequency could be based on a
sort of average error over all the harmonics, and one or two dropping
out (even the fundamental) wouldn't matter. I haven't actually tried
this one -- as usual, other things came up.

But I don't suppose you're interested in these ideas, since they are all
based on a frequency model of neural signals. A different train of
thought (although the same train of impulses!).

I've got to run for a train, but I do like these ideas -- and it's important
to have all sorts of notions out there for consideration. I'll try and think
about it.

Peter