PCT vs IT: approaching a resolution

[From Bill Powers (930408.0700)]

Here are some more answers as to why the IT/PCT argument is
important.

Information theory -- that is, the theory of signal transmission
-- must have some real applications in a brain, which is a signal
transmission device. But for some reason, when IT advocates try
to apply IT to PCT, the resulting language carries connotations
that seem opposed to the principles of PCT.

The main example of this has been language that suggests that
information from the disturbance gets into the control system and
serves as the basis for constructing an output signal to oppose
the disturbance. Even the information theorists recognize (more
or less) the paradox here: if the output opposes the disturbance,
it destroys information that is needed for constructing the
proper output. It is remarked that control, therefore, must be
"imperfect" -- but the paradoxical nature of that imperfection
has not been brought out sufficiently. The output or effect of
output, in fact, resembles the disturbance more closely as less
information is allowed to pass into the system. There is
something wrong with the idea that information in the disturbance
is used to construct the output -- that is, information in the IT
sense, log(R/r), which is directly related to uncertainty.

I am convinced that this problem, along with others we have had,
arises not from information theory itself, but from the way it is
being linked to phenomena. If we can resolve this problem, then
information theorists will be free to apply IT to problems of
signal transmission without appearing to deny the basic concepts
of PCT, and without paradoxes.

On Martin Taylor's advice, I looked up early references on
information theory. Not having the Bible on hand, I looked at my
old 1958 edition of the Britannica. The article on IT, as it
turned out, was written by Claude E. Shannon. One interesting
aspect of this article is how little the basic strategies of
thought have changed in the last 35 years. Shannon even talks
about filtering and prediction, and about cryptography and
linguistics, subjects that have been broached on this net in a
way that made me think the ideas were new.

The most interesting idea that struck me was that information
theorists on CSGnet seem to be treating disturbances as if they
were noise in a channel being used for communication. Everything
they say fits this interpretation as I found it in the article.
This would explain many of the baffling disputes among the
parties to this discussion.

If you interpret an external disturbance as a _message_, all of
the confusion drops away. The message contained in the variations
of the external disturbance is passed through a receiver, the
sensory receptors, and is expressed as modulation of a new
carrier, a neural signal. The actual noise in this channel is not
the disturbance, but thermal and chemical noise in the receptors,
and the noise inherent in pulse-frequency modulation. That noise
occuppies a completely different bandwidth from the variations in
the disturbance, and is of a far lower amplitude.

So we are no longer concerned with the disturbance as a noise
source that interferes with the operation of the control system.
It now becomes a smoothly-varying signal carried by a slightly
noisy channel. The channel noise, in fact, becomes only a small
fraction of the total variations that the disturbance alone would
cause in the perceptual signal.

As soon as we begin to treat the disturbance as a systematic
message instead of noise, we are out of the realm of information
theory. What happens to the message and what it means are of no
direct concern to IT. But the message, and the way it is handled
by the control system, is the central concern of PCT. IT can tell
us the limits on the fidelity with which this message can be
carried from one place to another -- it tells us, for example,
that the message can be reproduced only up to a certain rate of
variation, at which point channel noise and bandwidth limitations
will begin rendering it uncertain. But as long as the variations
in the message remain in the "safe" region, we can simply deal
with the message and forget about the noise.

The message, of course, is the amplitude of the disturbance as a
function of time. The message exists originally as an analog
quantity, and it appears, slightly degraded and in a different
physical form, as an analog quantity in the control system. It
does not appear alone, of course; to it is added another message
(when the loop is closed) about the amplitude of the output or
output effects as a function of time. When these two analog
representations are added together, the sum can no longer be
separated into the individual contributions.

If the two contributions are opposed, only the difference between
the amplitudes appears as a message carried by the perceptual
signal. When this difference becomes small enough, channel noise
becomes significant with respect to the remaining message, and
the remaining message becomes uncertain. The smaller the
difference, the greater the relative uncertainty. This
uncertainty puts the ultimate limit on the accuracy with which
the effects of output can cancel the effects of the disturbing
variable.

This is now a very different situation from the one we have been
arguing about. The uncertainty in the perceptual signal (with no
feedback) is not the whole amplitude of the disturbance as the IT
analysts have been assuming, but only the extent to which the
perceptual signal fails to be an accurate representation of the
disturbance. Probabilistic calculations do not apply to the whole
effect of the disturbance, but only to the slight deviations of
the perceptual signal from being a perfect representation of the
disturbance.

When the loop is closed, the amplitude of the perceptual signal
variations due to the disturbance is greatly reduced. But the
remaining amplitude is still much greater than the channel noise
until it becomes only a few percent of the amplitude of the
opposing signals. At that point, and only then, we begin to see
truly random variations in the perceptual signal, variations that
are not correlated with the disturbance or the output or anything
else we can notice. This is the realm of information theory and
probability calculations.

The "imperfections" in a control system due to noise appear only
when perfection has been approached so closely that channel noise
prevents a closer approach. If one thinks of the disturbance
itself as noise, then the ENTIRE signal is noise, and it becomes
hard to see how control could work at all. But if we see the
signals in the control loop as messages carried by a slightly
noisy channel, the paradox disappears; most of the error
correction takes place in a signal amplitude region where noise
in the channel is negligible. It now makes sense that a control
system could be 99% perfect even in the presence of disturbance
effects -- not noise -- that have an amplitude greater than that
of the largest possible perceptual signal.

Disturbing variables in the environment are not white noise or
any other kind of noise. They are physical variables affected in
systematic ways by other physical variables according to regular
laws, many of which we understand. Their effects on controlled
variables or perceptual signals are not random; they are highly
systematic and are, at least for the lowest level of perception,
well-understood. The amount of statistical uncertainty involved
in the behavior of disturbances and the consequent effects on
perception is normally only a small fraction of the magnitude of
the systematic effect. We are dealing with large signal
variations that have only a small amount of noise riding on them.

When we do simulations of control systems, we are dealing with
quantities and signals that have almost no noise in them at all
-- the only noise is in the rounding errors of calculation. This
is why a simple control-system model can approach perfect control
within a few parts per million.

But even in simulations, there are limits on bandwidth and speed.
These have been confused, in our discussion, with the limits
imposed by information theory. These limits, however, are not set
by uncertainty at all. They are set by systematic phase,
frequency, and amplitude relationships that are continuous in
nature, quantitative, and not the least bit uncertain. If you put
a mass in the controlled variable, and use a control system with
a single integrator in it, the ensemble will be unstable and will
oscillate. The oscillation is highly systematic. The instability
is not random. Information theory has nothing to do with this
phenomenon, this kind of imperfection of control. Reducing the
noise in the system or increasing the bandwidth of the control
system's components will, if anything, make the instabilities
worse. We are talking here about relationships among signals and
variables, not between signals and variables on the one hand, and
noise on the other hand.

I think this brings us close to a correct understanding of the
role of information theory in PCT. I don't think that the message
will be welcome. ITers have been trying to make information
theory, which is fundamentally a probabilistic approach, handle
phenomena which are actually perfectly regular and lawful, using
the techniques appropriate to situations where lawful
relationships are absent. If the application of IT is limited to
those areas where random phenomena actually occur, its scope will
become very much more restricted. I do not believe that any
important aspects of organized behavior involve significant
random processes -- except reorganization itself.

···

-------------------------------------------------------------
Best,

Bill P.

[Martin Taylor 930415 19:40]
(Bill Powers 930408.0700)

for some reason, when IT advocates try
to apply IT to PCT, the resulting language carries connotations
that seem opposed to the principles of PCT.

This statement is obviously true, at least for some readers. I have
hoped to write with sufficient skill to ensure that those connotations
are suppressed, but without success. I hope these connotations have
not affected all my readers.

The main example of this has been language that suggests that
information from the disturbance gets into the control system and
serves as the basis for constructing an output signal to oppose
the disturbance.

(Puzzlement here).

The most interesting idea that struck me was that information
theorists on CSGnet seem to be treating disturbances as if they
were noise in a channel being used for communication.

I never thought of it this way, and if that was something you got out
of my writing, it was because the idea never occurred to me, to counteract
it.

If you interpret an external disturbance as a _message_, all of
the confusion drops away. The message contained in the variations
of the external disturbance is passed through a receiver, the
sensory receptors, and is expressed as modulation of a new
carrier, a neural signal. The actual noise in this channel is not
the disturbance, but thermal and chemical noise in the receptors,
and the noise inherent in pulse-frequency modulation. That noise
occuppies a completely different bandwidth from the variations in
the disturbance, and is of a far lower amplitude.

That's more the way I have thought of it, though the word "message"
also carries unwanted freight.

The channel noise, in fact, becomes only a small
fraction of the total variations that the disturbance alone would
cause in the perceptual signal.

Yes, indeed. At least when control is good.

As soon as we begin to treat the disturbance as a systematic
message instead of noise, we are out of the realm of information
theory.

NO.

IT can tell
us the limits on the fidelity with which this message can be
carried from one place to another -- it tells us, for example,
that the message can be reproduced only up to a certain rate of
variation, at which point channel noise and bandwidth limitations
will begin rendering it uncertain. But as long as the variations
in the message remain in the "safe" region, we can simply deal
with the message and forget about the noise.

Up to a point. But always, the limit to the effectiveness of control
will be set by how well the perceptual system can resolve the CEV.

If the two contributions are opposed, only the difference between
the amplitudes appears as a message carried by the perceptual
signal. When this difference becomes small enough, channel noise
becomes significant with respect to the remaining message, and
the remaining message becomes uncertain. The smaller the
difference, the greater the relative uncertainty. This
uncertainty puts the ultimate limit on the accuracy with which
the effects of output can cancel the effects of the disturbing
variable.

Yes, that's what I have been trying to get at. This is well said.

This is now a very different situation from the one we have been
arguing about.

What YOU have been arguing about. Not me.

The uncertainty in the perceptual signal (with no
feedback) is not the whole amplitude of the disturbance as the IT
analysts have been assuming, but only the extent to which the
perceptual signal fails to be an accurate representation of the
disturbance.

Nope. The uncertainty in the perceptual signal is determined by the
resolution of the perceptual apparatus. It has nothing to do with the
disturbance. If the control were cut off, so that there was no output,
then the uncertainty in the perceptual signal would be largely determined
by the amplitude of the disturbance. I've tried over and over to point
out that it is only the uncontrolled fraction of the disturbance that
shows up as such in the perceptual signal. We are dealing with a
control LOOP here.

Probabilistic calculations do not apply to the whole
effect of the disturbance, but only to the slight deviations of
the perceptual signal from being a perfect representation of the
disturbance.

They apply anywhere, but some applications are more useful than others.
It turns out to be useful for an Engineer/Designer to do the calculations
for the unopposed disturbance.

When the loop is closed, the amplitude of the perceptual signal
variations due to the disturbance is greatly reduced. But the
remaining amplitude is still much greater than the channel noise
until it becomes only a few percent of the amplitude of the
opposing signals.

This is a bone of contention, I think.

There are two fundamental reasons why control is not perfect. One is
in the control system dynamics; there may be insufficient low-frequency
gain to oppose a persistent disturbing influence, or the disturbance
may be continually increasing, or the available output power may be
insufficient for the disturbance, or some such. The other is the perceptual
inability to determine what needs to be corrected, and in an effective
control system, this is the major limit, in my view.

At that point, and only then, we begin to see
truly random variations in the perceptual signal, variations that
are not correlated with the disturbance or the output or anything
else we can notice. This is the realm of information theory and
probability calculations.

The "imperfections" in a control system due to noise appear only
when perfection has been approached so closely that channel noise
prevents a closer approach.

Most of the time, in other words, if you include over-fast changes in
the disturbance or in the relation between the output and the CEV, both
of which are accounted for in the information analysis.

If one thinks of the disturbance
itself as noise, then the ENTIRE signal is noise, and it becomes
hard to see how control could work at all.

I agree.

But if we see the
signals in the control loop as messages carried by a slightly
noisy channel, the paradox disappears; most of the error
correction takes place in a signal amplitude region where noise
in the channel is negligible. It now makes sense that a control
system could be 99% perfect even in the presence of disturbance
effects -- not noise -- that have an amplitude greater than that
of the largest possible perceptual signal.

The dynamic error correction occurs in this range, but only while
detectable error remains in the perceptual signal. As we discussed
many moons ago, there is no reason why perceptual signals that are
down in the noise should affect the output, since the "control" at that
level is simply wasted energy. It makes much more sense for the
comparator or the output function to have a dead zone around zero signal,
and once again, the information (actually the signal detection) analysis
shows how such dead zones fall naturally out of the signal statistics.

Disturbing variables in the environment are not white noise or
any other kind of noise. They are physical variables affected in
systematic ways by other physical variables according to regular
laws, many of which we understand. Their effects on controlled
variables or perceptual signals are not random; they are highly
systematic and are, at least for the lowest level of perception,
well-understood.

Careful about viewpoint here. The theorist understands, but is that
understanding built into the PIF, or provided to it by an imagination
function in the ECS? If it is, then the ECS would not treat it as
white noise, and its equivalent rectangular bandwidth (or information
rate) would be accordinagly reduced.

When we do simulations of control systems, we are dealing with
quantities and signals that have almost no noise in them at all
-- the only noise is in the rounding errors of calculation. This
is why a simple control-system model can approach perfect control
within a few parts per million.

Fair enough, but it won't be as perfect if the dynamics are wrong. And
simulations using megahertz bandwidth and picovolt resolutions won't
be as good simulations of human performance as will ones using realistic
limitations on resolution and bandwidth.

But even in simulations, there are limits on bandwidth and speed.
These have been confused, in our discussion, with the limits
imposed by information theory. These limits, however, are not set
by uncertainty at all. They are set by systematic phase,
frequency, and amplitude relationships that are continuous in
nature, quantitative, and not the least bit uncertain.

I have not been among those making this confusion. The IT limitations
are just that--limitations. If there is poor design, things will be
worse, but good design cannot exceed the IT limitations.

ITers have been trying to make information
theory, which is fundamentally a probabilistic approach, handle
phenomena which are actually perfectly regular and lawful, using
the techniques appropriate to situations where lawful
relationships are absent.

First half of this sentence is OK, but the second is not. The techniques
deal with the degree of lawfulness, not the binary concept of lawfulness
versus randomness.

I do not believe that any
important aspects of organized behavior involve significant
random processes -- except reorganization itself.

I'm not sure how to react to this comment. We accept randomness in
reorganization precisely because information is lacking that could
guide it. We accept well structured control systems precisely because
the world in which they work is lawful enough ("contains" information)
to allow the structures to remain effective for periods longer than
it takes to build them. Both are readily understood in the same IT
framework. Bill's comment here reminds me of someone who might say
"gravity is all right where it belongs, in explaining planetary orbits,
but don't try to convince me it works here on Earth."

The IT argument is a structural one. Why are control systems as they are?
IT does not deal with specific dynamics of particular control systems,
though the analyses might well be applicable. It deals with HOW they
work and why the fundamental laws of control systems function in a real,
partially lawful, world. It's a bit funny, really, to be told that
PCT is necessary because the organism knows nothing about the influences
that act on its perceptions (and its health), and on the other hand be
told:

Disturbing variables in the environment are not white noise or
any other kind of noise. They are physical variables affected in
systematic ways by other physical variables according to regular
laws, many of which we understand. Their effects on controlled
variables or perceptual signals are not random; they are highly
systematic and are, at least for the lowest level of perception,
well-understood.

It is because this is true that control can work. It is because this
is true that an IT analysis can show why they work. And that is what
we have begun to do.

Martin

[Allan Randall (930420.1100 EDT)]

Bill Powers (930408.0700)

... Even the information theorists recognize (more
or less) the paradox here: if the output opposes the disturbance,
it destroys information that is needed for constructing the
proper output. It is remarked that control, therefore, must be
"imperfect" -- but the paradoxical nature of that imperfection
has not been brought out sufficiently.

This "paradox" is just the notion of controlling via detection of
error. You recognize this when you call the output of the comparator
an "error" signal. A perfect error controller has no error, yet
it achieves this by computing error. Obviously, perfect control can
only be approached. If this is a paradox, you will have to deal with
it, instead of sweeping it under the rug by invoking nonexistant
infinities.

The most interesting idea that struck me was that information
theorists on CSGnet seem to be treating disturbances as if they
were noise in a channel being used for communication.

I am really not sure what was said that gave you this idea. Your
misunderstanding of us could not be more complete. I *have* been
talking about disturbance as a message - not as noise, although it
can perhaps be viewed that way. Ashby argued for the existence of two
channels: one channel goes from reference to percept, the other
goes from disturbance to percept. The first interpretation defines
the disturbance as a message (which is the context we have been
dealing with mostly so far). The second interpretation defines
disturbance as noise. So far, we have been usually assuming a fixed
reference, so it is the disturbance as message approach we have
been dealing with in this debate.

As soon as we begin to treat the disturbance as a systematic
message instead of noise, we are out of the realm of information
theory.

Huh? I must confess this really threw me. Your whole article echoes
back many of the informaton theoretic ideas Martin and I have been
arguing for. Information theory is all about what makes a message
"systematic" instead of random. The possibility of systematic
disturbances instead of random ones is exactly what puts this whole
discussion *within* the realm of information theory. Ashby's Law
is exactly about taking advantage of such systematicity in order to
control. Ashby's Law would be unneccessary (not false, just not of
much use) if disturbances were all random.

What is this systematicity that you talk about? Try to define it in
a rigorous mathematical way. You contrast systematicity with
randomess. Try to state what you mean by this without introducing
the notion of probabilities. You may just end up reinventing
information theory.

I want to make one thing VERY clear. The "information theory" you
argue against in this article has absolutely nothing whatsoever
to do with real information theory. You have to stop thinking that
information is simply the capacity of a channel, or applies only to
random systems. Complete randomness is one end of the information
spectrum. There would be no point in measuring entropy if maximum
entropy was the only result possible. Your own arguments (which you
seem to think counter to information theory) about systematic,
nonrandom messages from the disturbance sounds VERY much like exactly
what Martin and I have been saying all along! It sounds to me like you
just have a problem with the term information. If you decide to agree
with us, except you wish to call it "systematicity" instead of
"information," then I have no problem with that. That is a rather
minor quibble over terminology, and not a conceptual disagreement.

The message, of course, is the amplitude of the disturbance as a
function of time. The message exists originally as an analog
quantity, and it appears, slightly degraded and in a different
physical form, as an analog quantity in the control system.

Correct me if I'm wrong. You seem to be saying that there is a
message from the disturbance in the percept. This message is only
"slightly degraded," and is combined with another message when
there is a closed loop. If so, could you please tell me the
difference between this and saying that there is information from
the disturbance in the percept?

If you respond by saying that the disturbance message is added to
the output message, and so *all* information is destroyed, please
tell me why it makes sense for you to say that the disturbance
message is only "slightly degraded." Complete destruction of a
message is pretty total degradation, I would say.

Disturbing variables in the environment are not white noise or
any other kind of noise ... Their effects on controlled
variables or perceptual signals are not random; they are highly
systematic ...

Information.

···

-----------------------------------------
Allan Randall, randall@dciem.dciem.dnd.ca
NTT Systems, Inc.
Toronto, ON