[From Bill Powers (930408.0700)]
Here are some more answers as to why the IT/PCT argument is
important.
Information theory -- that is, the theory of signal transmission
-- must have some real applications in a brain, which is a signal
transmission device. But for some reason, when IT advocates try
to apply IT to PCT, the resulting language carries connotations
that seem opposed to the principles of PCT.
The main example of this has been language that suggests that
information from the disturbance gets into the control system and
serves as the basis for constructing an output signal to oppose
the disturbance. Even the information theorists recognize (more
or less) the paradox here: if the output opposes the disturbance,
it destroys information that is needed for constructing the
proper output. It is remarked that control, therefore, must be
"imperfect" -- but the paradoxical nature of that imperfection
has not been brought out sufficiently. The output or effect of
output, in fact, resembles the disturbance more closely as less
information is allowed to pass into the system. There is
something wrong with the idea that information in the disturbance
is used to construct the output -- that is, information in the IT
sense, log(R/r), which is directly related to uncertainty.
I am convinced that this problem, along with others we have had,
arises not from information theory itself, but from the way it is
being linked to phenomena. If we can resolve this problem, then
information theorists will be free to apply IT to problems of
signal transmission without appearing to deny the basic concepts
of PCT, and without paradoxes.
On Martin Taylor's advice, I looked up early references on
information theory. Not having the Bible on hand, I looked at my
old 1958 edition of the Britannica. The article on IT, as it
turned out, was written by Claude E. Shannon. One interesting
aspect of this article is how little the basic strategies of
thought have changed in the last 35 years. Shannon even talks
about filtering and prediction, and about cryptography and
linguistics, subjects that have been broached on this net in a
way that made me think the ideas were new.
The most interesting idea that struck me was that information
theorists on CSGnet seem to be treating disturbances as if they
were noise in a channel being used for communication. Everything
they say fits this interpretation as I found it in the article.
This would explain many of the baffling disputes among the
parties to this discussion.
If you interpret an external disturbance as a _message_, all of
the confusion drops away. The message contained in the variations
of the external disturbance is passed through a receiver, the
sensory receptors, and is expressed as modulation of a new
carrier, a neural signal. The actual noise in this channel is not
the disturbance, but thermal and chemical noise in the receptors,
and the noise inherent in pulse-frequency modulation. That noise
occuppies a completely different bandwidth from the variations in
the disturbance, and is of a far lower amplitude.
So we are no longer concerned with the disturbance as a noise
source that interferes with the operation of the control system.
It now becomes a smoothly-varying signal carried by a slightly
noisy channel. The channel noise, in fact, becomes only a small
fraction of the total variations that the disturbance alone would
cause in the perceptual signal.
As soon as we begin to treat the disturbance as a systematic
message instead of noise, we are out of the realm of information
theory. What happens to the message and what it means are of no
direct concern to IT. But the message, and the way it is handled
by the control system, is the central concern of PCT. IT can tell
us the limits on the fidelity with which this message can be
carried from one place to another -- it tells us, for example,
that the message can be reproduced only up to a certain rate of
variation, at which point channel noise and bandwidth limitations
will begin rendering it uncertain. But as long as the variations
in the message remain in the "safe" region, we can simply deal
with the message and forget about the noise.
The message, of course, is the amplitude of the disturbance as a
function of time. The message exists originally as an analog
quantity, and it appears, slightly degraded and in a different
physical form, as an analog quantity in the control system. It
does not appear alone, of course; to it is added another message
(when the loop is closed) about the amplitude of the output or
output effects as a function of time. When these two analog
representations are added together, the sum can no longer be
separated into the individual contributions.
If the two contributions are opposed, only the difference between
the amplitudes appears as a message carried by the perceptual
signal. When this difference becomes small enough, channel noise
becomes significant with respect to the remaining message, and
the remaining message becomes uncertain. The smaller the
difference, the greater the relative uncertainty. This
uncertainty puts the ultimate limit on the accuracy with which
the effects of output can cancel the effects of the disturbing
variable.
This is now a very different situation from the one we have been
arguing about. The uncertainty in the perceptual signal (with no
feedback) is not the whole amplitude of the disturbance as the IT
analysts have been assuming, but only the extent to which the
perceptual signal fails to be an accurate representation of the
disturbance. Probabilistic calculations do not apply to the whole
effect of the disturbance, but only to the slight deviations of
the perceptual signal from being a perfect representation of the
disturbance.
When the loop is closed, the amplitude of the perceptual signal
variations due to the disturbance is greatly reduced. But the
remaining amplitude is still much greater than the channel noise
until it becomes only a few percent of the amplitude of the
opposing signals. At that point, and only then, we begin to see
truly random variations in the perceptual signal, variations that
are not correlated with the disturbance or the output or anything
else we can notice. This is the realm of information theory and
probability calculations.
The "imperfections" in a control system due to noise appear only
when perfection has been approached so closely that channel noise
prevents a closer approach. If one thinks of the disturbance
itself as noise, then the ENTIRE signal is noise, and it becomes
hard to see how control could work at all. But if we see the
signals in the control loop as messages carried by a slightly
noisy channel, the paradox disappears; most of the error
correction takes place in a signal amplitude region where noise
in the channel is negligible. It now makes sense that a control
system could be 99% perfect even in the presence of disturbance
effects -- not noise -- that have an amplitude greater than that
of the largest possible perceptual signal.
Disturbing variables in the environment are not white noise or
any other kind of noise. They are physical variables affected in
systematic ways by other physical variables according to regular
laws, many of which we understand. Their effects on controlled
variables or perceptual signals are not random; they are highly
systematic and are, at least for the lowest level of perception,
well-understood. The amount of statistical uncertainty involved
in the behavior of disturbances and the consequent effects on
perception is normally only a small fraction of the magnitude of
the systematic effect. We are dealing with large signal
variations that have only a small amount of noise riding on them.
When we do simulations of control systems, we are dealing with
quantities and signals that have almost no noise in them at all
-- the only noise is in the rounding errors of calculation. This
is why a simple control-system model can approach perfect control
within a few parts per million.
But even in simulations, there are limits on bandwidth and speed.
These have been confused, in our discussion, with the limits
imposed by information theory. These limits, however, are not set
by uncertainty at all. They are set by systematic phase,
frequency, and amplitude relationships that are continuous in
nature, quantitative, and not the least bit uncertain. If you put
a mass in the controlled variable, and use a control system with
a single integrator in it, the ensemble will be unstable and will
oscillate. The oscillation is highly systematic. The instability
is not random. Information theory has nothing to do with this
phenomenon, this kind of imperfection of control. Reducing the
noise in the system or increasing the bandwidth of the control
system's components will, if anything, make the instabilities
worse. We are talking here about relationships among signals and
variables, not between signals and variables on the one hand, and
noise on the other hand.
I think this brings us close to a correct understanding of the
role of information theory in PCT. I don't think that the message
will be welcome. ITers have been trying to make information
theory, which is fundamentally a probabilistic approach, handle
phenomena which are actually perfectly regular and lawful, using
the techniques appropriate to situations where lawful
relationships are absent. If the application of IT is limited to
those areas where random phenomena actually occur, its scope will
become very much more restricted. I do not believe that any
important aspects of organized behavior involve significant
random processes -- except reorganization itself.
···
-------------------------------------------------------------
Best,
Bill P.