[From: Bruce Nevin (Wed 93106 14:26:15 EDT)]
Well, I got caught in a cascade anyway, so here's a try at catching up.
Rick Marken (931005.0900) --
Thanks for examining your perceptions as I asked. I think this is a
summary of your findings.
[p`] [p] [b]
pin bin pin vs. bin
pud bud "pud" vs. bud
s[p`]in spin sbin all 3 = spin
I call "pud" a non-word, meaning a word-perception that happens not to be
in my lexicon, has no meaning. You say, I think equivalently,
It is a word (a sound event); it just has no meaning for me when I
just pay attention to its as a "word" event.
I claim that contrast is fundamental, and that what we think of phonemes
or letters or distinctive features are ways of representing the contrasts
between utterances or possibly between syllables. One cannot produce p
and b in isolation in order to perceive a relation of contrast between
them. One cannot produce voicelessness and voicing apart from producing
two utterances. The p and b sound-segments, or the sound features of
voicelessness and voicing, become perceptual means (together with other
perceptual means) for controlling the perception of contrast between
utterances. Utterances must be discrete and recognizable if we are to
communicate with them as we do. A child first learns a word as a fixed
sound sequence. Then the child learns to represent the contrasts between
words by a small set of phonemic elements--whether these are sound
segments, or sound features, or something else, is immaterial here, and
indeed it may well vary. It is important to recognize that this is (I
claim) a result of a perceptual control process that compares words that
are not the same, factors out elements by which they differ, and then
uses those factored elements (which represent the contrasts between
words) as means for recognizing which word is spoken and for pronouncing
words in a way that others recognize. In the social situation of
language use there is selective pressure, so to speak, for the behavioral
outputs (pronunciations of words) to be recognizable. Different
representations (different factorings into controllable perceptions)
might result in different individuals, without their pronunciations
(behavioral outputs) diverging so greatly as to be unrecognizeable. The
factored-out controllable perceptions (controlled perceptions) are the
phonemic elements or phonemes. The gain on control of a given phonemic
element at a given time is a function of the means that it serves: the
contrast of the given word with other words that it might be mistaken
for.
It happens that, in English, there are syllables with initial b that
might be mistaken for syllables with initial p (bin vs. pin). It also
happens that, in English, one can pronounce b or p or an intermediate
sound interchangeably between syllable-initial s and a vowel. One can
pronounce either p or b there, and (there being no contrast) one normally
pronounces an intermediate sound. The sound-features or sound-segments
factored out as a phoneme p or whatever for pin vs bin are not
fundamental. What is fundamental is contrast, and the phonemic elements
are perceptual means for controlling that more fundamental perception of
contrast. There is no perception of contrast between voiced and
voiceless-aspirated elements after syllable-initial s, so those
perceptual means are never used there. For a speaker of English, they do
not exist there. The gain on control of the controlled perception by
means of which I normally make damned sure I say pin instead of bin, or
vice versa, drops to zero when I say spin.
···
=========================
Rick, I did say to Bill "you give up too easily" but later on in the same
message I agreed with him that "nobody knows why ... people perceive a
contrast ... between pin and bin, but not between spin and sbin, ...,
not even a linguist". The details behind a given state of a language are
historically contingent in the ongoing process of individuals creating
and maintaining the language as a product of social use.
But reading on I see you noticed that I'm not at all claiming that I know
the explanation. But it wasn't a "bluff". (What category perceptions
are you laboring with, anyway?) I was objecting to merely throwing up
our hands and saying it's unknowable. A great deal is in fact known
about particular cases, and more is being learned daily. The knowledge I
think is of two sorts:
Historical: A documented and/or reconstructed prior state of the language
was thus, and the process of change from x at that stage to y at present
appears to have thus and such an explanation. Labov's work on the social
motivation of a sound change on Martha's Vineyard is a good instance of
this sort of account.
Generic or universal: (a) The possible states that languages can be in are
constrained by thus and such factors, or (b) thus and such characteristics
appear to be found in all known languages and so there must be some
biological or physical or similar constraint that accounts for it, we
just don't know it yet. Most universals of the latter sort are
"tendencies" that admit of exceptions. PCT I think promises to put (b)
kinds of accounts on a scientific footing for the first time by
converting them to (a) kinds of accounts.
Martin Taylor 931005 13:30 --
Thanks for the effort at clarification. I confess to some long-standing
puzzlement about event perceptions, which may as well come out here as a
digression from the main thread. As I understand it an event is a
sequence that differs from a sequence perception in that it is short,
familiar, and well practiced (bit of redundancy there). I have trouble
conceiving of an event perception as continuous, as opposed to the
discreteness of category perceptions. But now I recall the discussion we
had some while ago about the curious shift in "speed" of perceptions in
crisis, and your anecdote about diving for a cricket ball, which Bill
suggested was due to shifting attention to the event level. Is the
apparent liesureliness of things that actually are transpiring very
rapidly due to dropping from the discrete states of the category level to
continua on the event level? Somebody help me out here.
Another difficulty, which I have mentioned before. Phonemes appear to be
category perceptions, as you have said. But words are characterized as
event perceptions in BCP. Granted, this was when Bill was conceiving of
phonemes as sounds rather than categories of sounds (or of
sound-features). Words must be short, familiar sequences of categories.
Whereas most category perceptions are associated with words, phoneme
perceptions are not (except in the science and philosophy sublanguages
developed for linguistics). Some short, familiar sequences of words are
treated as units with meaning not associated with the constituent words
(idioms, frozen expressions, cliches), other sequences of words
constitute syntactic constructions of varying complexity. There seems to
be some looping within the category and sequence levels of the sort that
you suggest in your last paragraph.
Rick Marken (931005.1300) --
I have many perceptions that I would
be comfortable calling "contrasts"; my perception of the relationship
between bin and pin is one, as well as my perception of the relationship
between the height of a very short and a very tall person, etc.
The former is categorial, binary, either/or; the latter is scalar,
relative, continous. Seven feet is tall for a man, short for a giraffe.
There are men of whom you say they are of middle height, neither tall nor
short, but your perceptual control hierarchy compels you and Bill
(evidently) to assign a sound that is demonstrably intermediate between b
and p to the p phoneme as against the b phoneme. Martin has emphasized
the same point about categorial vs. sub-categorial perceptions.
I think that limits a nice
word a bit (why not include perceived difference between anything -- like
short and tall people)
This is a technical definition of the word, analogous to the restrictions
that are imposed in the technical definition of "control" in PCT.
Substitute "phonemic contrast" wherever I have used "contrast" with
reference to phonemes, etc. Perhaps that will help. The reason it is
special for language is that the pronunciation of a word must be
recognizeable and distinguished from pronunciations of other words
(including possible words not actually in use, such as pud). See above
for more.
Martin Taylor 931005 16:00 --
I am saying a bit more about contrast as a perception that is more
fundamental than the phonemic-element perceptions, at least in
developmental terms (the child language learner), and I think also in
general (varying the gain depending on the informational load). I don't
feel that I have quite got a coherent sense of it, and what I say is
doubtless not as coherent as it ought to be, but I'm going to be
tenacious about the intuition behind it until it does come clear. If you
can see what I'm driving at above, I would appreciate any help you can
give. I don't see contrast in this fundamental sense as a perceived
relation between phonemes, but rather as a perceived relation between
utterances (words or syllables) constituted of/factored into phonemes,
such that the phonemic elements are perceptual means for controlling the
perception of utterance contrast.
By the way, I think I'm going to have a hard time avoiding making mine a
PCT dissertation, so these issues will assume more and more pressing
relevance for me as time goes on.
PEPINSKY Hal, Criminal Justice Tue, 5 Oct 1993 12:53:27 CDT --
Sbin is more complicated to pronounce than spin. S, like p, is
unvocalized. To pronounce sbin, you have to time your vocal chords to
go into action after the s and before the b. It can be done, but not
as readily as zbin or spin.
When you say "mouse bin" do you experience this difficulty?
Starting up voicing is physiologically independent of moving the tongue
and lips. However, it does require more effort to start voicing while
the mouth (and nasal passage) is closed, because air pressure must be
higher below the larynx than above it in order to sustain voicing. Bill
observed this. Therefore it is harder to say bin than it is to say
either in or pin.
Here's a crude graphical representation:
"pin"
Lips p
Tongue nnn
Larynx hiiiinnn
Velum nnn
"bin"
Lips bb
Tongue nnn
Larynx biiiinnn
Velum nnn
"spin"
Lips pp
Tongue ss nnn
Larynx iiiinnn
Velum nnn
Time is on the horizontal dimension. Letters aligned with the name of a
particular articulator (Lips, Tongue, Larynx, Velum) indicate closure or
the relevant proximal closure for that articulator at the indicated
stretch of the time dimension. For example, during the n segment,
simultaneously: the tongue closes the mouth, the larynx is vibrating, and
the velum is lowered so as to open the nasal passage for nasal resonance.
PEPINSKY Hal, Criminal Justice Tue, 5 Oct 1993 15:38:00 CDT --
And yes, you can feel the vibration of voicing by touching your larynx.
I wasn't sure that the relative timing could be perceived that way in a
convincing way.
Bill Powers (931006.0745 MDT) --
This is excellent. A return to Ma Reality for some grounding.
pin and bin result in separate perceptual signals, but
spin and sbin result in a signal from only one perceptual
function because there is only one that responds to either spin
or sbin (slightly less to sbin).
I suspect that you're contrasting sbin with "spin" as ordinarily
pronounced. The aspirated term should sound and feel something like
sp-hin, with what feels like effortful puffing of an h after the p, quite
different from the normal pronunciation of spin.
Unfortunately, when I vary the voice
pitch while holding the vowel constant, there are also radical
shifts in the fine structure, so this is still not the right
approach. Now there is TOO MUCH sensitivity to differences,
particularly to differences that aren't supposed to matter.
Perhaps the frequency information contained in the frequency-
tracking system could be used to insert compensating effects in
the waveform-discriminating system. But that begins to sound
kludgy.
Could you normalize pitch to a constant? The system accomplishing this
would simultaneously provide control of a perception of pitch contours
for intonation patterns, one of the first things infants learn, according
to some investigators at least.
The correlellogram business was in a talk announcement that I forwarded
to the net. I have no information other than that in the announcement,
sorry.
There is an article in the latest issue of _Language_ (the Journal of the
Linguistic Society of America) to which I would like to direct your attention:
Johnson, Keith, Edward Flemming, & Richard Wright.
The hyperspace effect: phonetic targets are hyperarticulated.
_Language_ 69.3:505-528 (September 1993).
They report some experiments in which subjects choose from a palatte of
synthesized vowel sounds the vowel that best fits (a) a heard vowel or
(b) their own pronunciation of a given word. They show that variation
within a given speaker, and also variation across speakers (all the same
southern California dialect) is lots less than the variation that puzzled
them, and whose explanation is the subject of the paper, to wit: they
consistently chose vowels that might serve as targets for
"hyperarticulated" pronunciations, pronunciations exaggerating the
phonemic contrasts between words. This seems to bear pretty directly on
the issues that concern us here. I have not had time to finish it yet,
maybe tonight on the subway.
They refer to "Lindblom's (1990) `Hyper- and Hypoarticulation' (H&H)
theory of phonetic variation".
Lindblom 1990 conceptualizes speech production as a feedback system
in which the input is the goal and the actual production is the output.
The extent to which the output matches the input goal depends on the
gain, or amplification, of the feedback loop. The gain of the feedback
loop is thus analogous to effort. The salient feature of this model is
that the input goal is the most distinctive, hyperarticulated speech,
since this is the signal that the output approximates as the gain is
maximized. Reduction processes are not conceived of as altering the
goals, but result from expending less effort and thus falling short of
the goal.
The reference:
Lindblom, Bjo"rn. 1990. Explaining phonetic variation: A sketch of the
H&H theory. Speech production and speech modelling, ed. by William J.
Hardcastle and A. Marchal, 403-439. Dordrecht: Kluwer.
It was this that suggested to me that it might just be politically
possible to put across a PCT dissertation, if this sort of thing is out
there in the literature so that I can lean on it in citations.
Gotta run. This has been fun but I have work to do. And after all this
I would like to stay employed! (Oops, maybe I didn't say: looks like I
can finagle a commuting route to the new location, wherever it settles
down to be, using two trains and carpooling with a coworker from
Winchester to the job. Fingers crossed!)
Bruce
bn@bbn.com