PCT & IT

[From Bill Powers (930416.0800 MDT)]

Martin Taylor (930415.1400) --

Otherwise you will leave out the high-frequency variations that
are beyond the arbitrary limit of the rectangular bandwidth.

Of course. That's obvious. But note the word "equivalent." I
use the general approach of Blackman and Tukey (The measurement
of power spectra, p24 in the 1958 Dover edition) in dealing with
the problems of fractional degrees of freedom.

What does that have to do with reconstructing a waveform? A power
spectrum is of as little use as a rectangular distribution for
reconstructing a waveform. What you need are the specific
amplitude and phase values for every frequency out to the limit
where there is enough amplitude left to make a difference in the
waveform. A power spectrum throws away sign information and phase
information, both of which are essential for the reconstruction.

I did not say that a signal is contained in a finite band. I
said that if it has finite power, it has an equivalent
rectangular bandwidth.

That is true, but irrelevant to reconstructing a waveform. The
"equivalent" bandwidth is equivalent only in those respects you
preserve in constructing it -- which does not include amplitude,
phase, or frequency. The equivalent rectangular bandwidth is an
artifact constructed for convenience in doing certain kinds of
calculations. It is not equivalent to the actual situation in any
of the ways that would be helpful in reconstructing a waveform.

Once more unto the quantized breach, dear friends. Let's get it
straight. I do not now, and never have, belonged to the
quantized party of America. I don't plead the Fifth. I don't
have friends who are quantized. They all behave continuously,
like good citizens.

Then why do you keep talking about D/r? Why not just let r go to
the limit of zero (the continuous case) and the uncertainty to go
infinity (log(D/r), and use those continuous equations? It seems
to me that all the basic definitions require that D/r be a finite
whole number, and that the boundary between d and d+r be fixed,
with no values of d between d and d+r. To me that is quantization
-- the division of a continuous scale into fixed intervals.

Bill points out that this [derivation of a relationship between
bandwidth and D/r) leads to absurd conclusions. In mitigation,
I can only quote my comment that went along with this equation
in my 930405 16:00 posting:

This is the maximum effective gain that can be achieved in any
control system, if my ever-reliable :frowning: algebra is working.

It obviously wasn't working, and that's where the argument
should have hit.

And what about your patronizing exasperation about my not paying
attention to the relationships you had laid out and that I had
not duly memorized?

The new equation is

G = (D/r)^(1-1/B) - 1

And you say about it

For a bandwidth ratio approaching unity, the maximum useful gain
approaches zero. All correct so far ---

That is to say, if the bandwidth of the disturbance is 1 hz, and
the bandwidth of the perceptual function is 1 Hz, the maximum
usable loop gain in the associated control system is zero. This
is simply not true. We will shortly have a new version of Simcon
that allows creating a random disturbance with a one, two or
three stage filter to permit varying the bandwidth and the
frequency rolloff. So you can create a disturbance with a
bandwidth of 1 Hz. Using an amplifier with a time constant of 1
sec (or whatever the right number would be) you can give the
control system an input bandwidth of 1 Hz. I suggest that before
you carry this line of deduction any further, you set up a
control system and test your "all correct so far" conclusion. The
mathematics may be all correct so far, but the conclusion is
still false. There is obviously a problem with the premises.

You will find that the control system can be adjusted to keep the
effects of the disturbance on the perceptual signal very small,
with a loop gain very much larger than 0. If you don't want to do
it, I will do it for you.

The bandwidth of the disturbance has absolutely nothing to do
with the maximum usable loop gain in a control system. There is
something wrong with your premises, which renders all the
manipulations that follow from them spurious.

... for small SNR one has to use the continuous information
formula, in which P and D now become the RMS values rather than
the ranges, and r is the RMS uncertainty of the sample after the
observation. And before we get into "the observation is an exact
number, whatever it is," remember that the uncertainty is that
of the CEV given the perceptual signal, not of the perceptual
signal given the perceptual signal.

In passing, I'm curious why there is such rejection on the IT
side of concepts like variance and correlation, which involve
exactly the same kinds of derived measures: "RMS" variations,
which is just the calculation of sigma. Correlation is simply
sigma(xy)/[sigma(x)*sigma(y)]. These measures are obtained from
discrete samples, and I'm sure that by rearranging the
expressions through equivalence transformations and zinging in a
logarithm here and there, you could show that the expressions are
closely related to those by which uncertainty is calculated
(conditional uncertainty in the case of sigma(xy)). The basic
premises are essentially identical.

As to "the uncertainty is that of the CEV given the perceptual
signal", that uncertainty has to be zero, because the CEV is
defined as the external correlate of the perceptual signal. You
are arguing from the premise that the CEV has an objective
existence independent of the control system, so that the
perceptual signal will represent it more or less accurately. Thus
the CEV could be in one state, while the perceptual signal
represents it as being in a somewhat different state. But that
contradicts the definition of a CEV.

The perceptual signal is the only true measure of the CEV. The
bandwidth of the perceptual signal IS the bandwidth of the CEV.

The valid formulae are
Perceptual information per sample = log ((P+r)/r)
Disturbance information per sample = log ((D+r)/r)

This can't be the correct set of formulae. The implication is
that the perceptual information per sample is independent of the
(imaginary) sampling frequency. When you let the sampling
frequency go to infinity to approach the continuous case, the
information per sample must decrease accordingly, or else the
information rate will go to infinity. It will go to infinity-
squared, because r also approaches zero.

This is very much like the problem of representing a continuous
control system on a digital computer. If you just naively compute
your way around the loop, you will end up with oscillations at
the frequency of iteration for any loop gain greater than 1. To
make the computations independent of the iteration frequency, and
get the right behavior, you must put in a slowing factor that
introduces physical time. As the iteration rate increases, the
amount of change permitted per iteration decreases, so the system
converges to the correct behavior as iteration frequency
increases without limit.

Your formulae above do not take physical time into account, which
is why they lead to an absurd result as the sampling frequency is
raised. The information per sample must decrease as sampling
frequency is increased. But there is no variable in the formulae
representing sampling frequency as samples per unit of physical
time. Without an explicit variable linking your discrete
calculations to physical time, you can't make the transition from
the discrete representation to the continuous (physical)
representation -- correctly.

There is no
reason that the perceptual bandwidth has to be greater than the
disturbance bandwidth.

The reason is that intrinsic delay, if you look from a
straightforward analogue signal-processing viewpoint.

I think you're confusing "disturbance bandwidth" with "the
highest frequency component of the disturbance." And note that a
pure delay doesn't change the bandwidth at all. Only an integral
lag (or some such) will cut off high frequencies. Integral lags,
by the way, introduce physical time.

And I think you're still equating "disturbance" to "change in the
CEV." You've never understood why I make the disturbance, or the
disturbing influence, independent of the state of the CEV. The
reason is precisely to avoid the sort of confusion we have here.
If the CEV is the position of a limb, the disturbance is not a
change in that position, but (for example) the magnitude of a
force applied to the limb. There is no necessary relationship
between the applied force (or its bandwidth) and the resulting
changes in the CEV, because there is at least one other force
being applied -- the output of the system controlling limb
position. And remember that the measure of the CEV is the
perceptual signal, by definition: the input function doesn't even
come into it. To deduce the hypothetical state of an external
equivalent of the CEV you would have to apply the inverse of the
perceptual function to the perceptual signal. You posted a
diagram yesterday that makes exactly that point.

The normal practice is to smooth the samples
until only the envelope is visible. Of what use would a 40 KHz
signal be at the terminals of a loudspeaker?

Yes, I did try to make that point, but you said it was
unnecessary, so I dropped it.

No, you have dropped it because you're now saying that the
bandwidth of the perceptual function has to be greater than that
of the disturbance. If the samples are smoothed, the bandwidth of
the perceptual function can be LESS than that of the disturbance.
I brought it up as an argument against the idea that B > 1 (as
then defined). If you now claim that point, it contradicts your
assertion that the perceptual bandwidth is greater than the
disturbance bandwidth (B > 1 as now defined).

If I said it was unneccessary, it was because the sampling is
imaginary in the first place.

But this is an S-R approach to the problem, something I have
been trying to avoid all along.

You can't avoid the S-R approach when you're talking about a
single function in a control system. All the functions are S-R
devices. In speaking of the effect of the disturbance on the
perceptual signal, it's perfectly legitimate to talk of its
contribution in the absence of output (which can be observed at
every zero-crossing of the output effect).

It should be easy enough to set up a simulation to test the
effect of increasing gain up to and then beyond the limit that
seems to be implied by the information analysis. (I take no
responsibility for the correctness of the algebra; it is the
principle that I stand behind.)

I agree about the simulation test, but come on! If you don't take
responsibility for the correctness of your algebra -- on which
all of your arguments depend -- who will? You've already gone off
on a number of deductive tangents as a result of making errors.
How can you convince anyone that your "principles" aren't equally
tainted by wrong deductions, when your defense of them depends
entirely on the mathematical manipulations you apply to your
premises, without any guarantee that your manipulations are free
of error? Or is your faith in IT so unshakeable that no mere
mathematical proof or disproof (or simulation) could disturb it?

As soon as we begin to treat the disturbance as a systematic
message instead of noise, we are out of the realm of
information theory.

NO.

YES. Why treat a continuous and regular phenomenon as if it is a
random variable? Why not use the simple and direct reasoning
appropriate to continuous variables when that is the kind of
phenomenon we are dealing with? Why not reserve the much more
complex and opinion-weighted and assumption-sensitive arguments
of uncertainty theory for situations in which the variables are
in fact unpredictable and irregular on the scale of interest?

Nope. The uncertainty in the perceptual signal is determined by
the resolution of the perceptual apparatus.

So it's premise against premise. I like my premise better.

The resolution of the perceptual apparatus is infinite. For any
range of the perceptual signal above the dead zone near zero and
below the saturation level near maximum, the frequency of the
neural signal can change by any arbitrarily small amount. It does
not jump from one finite frequency to another. Any change in the
inputs whatsoever is reflected as a corresponding change in the
frequency of the perceptual signal. There is no threshold amount
of input change required to produce a change in the perceptual
signal. There may be a least amount of change required for a
person to judge that a change has occurred (a much more complex
and higher-level process), but even these JNDs are not associated
with fixed intervals on the perceptual scale.

Thus the uncertainty in the perceptual signal is set by channel
noise, not by the resolution of the perceptual apparatus.

Probabilistic calculations do not apply to the whole
effect of the disturbance, but only to the slight deviations of
the perceptual signal from being a perfect representation of
the disturbance.

They apply anywhere, but some applications are more useful than
others. It turns out to be useful for an Engineer/Designer to do
the calculations for the unopposed disturbance.

You miss my point. If the disturbance is a pure sine wave, it
contains no uncertainty at all, although one can treat the sine
wave using the manipulations appropriate to random variables and
get numbers out of the manipulations. The numbers, however, are
meaningless, because there is in fact no uncertainty.

The only uncertainty that exists in the disturbance is the
_unpredictable_ component, which can range from almost none to
the entire signal. The mathematics of uncertainty should be
applied ONLY to the unpredictable component. This has nothing to
do with whether the disturbance is opposed or unopposed.

If the waveform of the perceptual signal is a perfect replica of
the disturbance, or if it can be expressed as a regular noiseless
function of the disturbance (as in a control-system model), then
there is no uncertainty involved. Uncertainty would enter only
when the function used to express the relation between the
perceptual signal and the disturbance fails to predict the
observed relationship; then the mismatch can be called an
uncertainty, and appropriately treated in terms of random
variables.

When the loop is closed, the amplitude of the perceptual signal
variations due to the disturbance is greatly reduced. But the
remaining amplitude is still much greater than the channel
noise until it becomes only a few percent of the amplitude of
the opposing signals.

This is a bone of contention, I think.

Not if we're talking about simulations. If a noise-free
simulation (to which the above remarks certainly apply) fits
behavior within a few percent, then my statement is probably true
of the real system as well.

There are two fundamental reasons why control is not perfect.
One is in the control system dynamics; there may be insufficient
low-frequency gain to oppose a persistent disturbing influence,
or the disturbance may be continually increasing, or the
available output power may be insufficient for the disturbance,
or some such. The other is the perceptual inability to
determine what needs to be corrected, and in an effective
control system, this is the major limit, in my view.

It is not the major limit in our simulations. It is hardly a
consideration at all. And in the behavior that is matched by our
noise-free models, it can't be the major factor. The variations
due to inability to determine what needs to be corrected, in our
tracking experiments, are exactly known: one pixel of change,
which is easily visible. The computed value of the reference
signal in these experiments, with slow disturbances, is usually a
fraction of a pixel away from zero. The actual range of the
variables is from 30 to 100 pixels, and the potential range (set
by VGA screen limits on my machine) is 480 pixels. The RMS
uncertainty between the model and the actual behavior (the
measure you mention above as appropriate for the continuous case)
amounts to 5% or less of the range of the variables.

If we lower the disturbance frequency enough, the person can keep
the cursor on the target with errors of only 1 or 2 pixels. So
the limit set by perceptual resolution (here enforced by the
screen resolution) is negligible with respect to the limits set
by dynamic considerations.

The IT argument is a structural one. Why are control systems as
they are? IT does not deal with specific dynamics of particular
control systems, though the analyses might well be applicable.
It deals with HOW they work and why the fundamental laws of
control systems function in a real, partially lawful, world.

You have yet to show that IT does in fact explain anything about
control systems. You have complained now and then that because
we're still hung up on the basics, you can't get on to the really
interesting stuff. But if you can't defend the basics, it's not
likely that the more interesting stuff will be interesting to
anyone but you.

It seems to me that the reasoning involved in information theory
depends to an inordinate degree on taking just the right attitude
toward it, making just the right interpretation (on which even
you and Allen don't seem to agree and which others versed in IT
like Cliff Joslyn also see differently). Reducing IT to practice
appears almost too difficult to do at all; just look at Allen's
proposed proof of something about entropy, which requires solving
an NP-hard (one might even say NP-impossible) problem on the way.
As a practical approach to explaining behavior, I am not
impressed with IT. I am not one of those who loves complexity for
its own sake. And I am definitely not impressed with the
manipulations I have seen so far, which are loaded with errors
and erroneous predictions. So far your principles have not done
well by you in terms of leading you to verifiable statements
about either real systems or simulations.

···

------------------------------------------------------------
Bill P.

[Martin Taylor 930416 2140]
(Bill Powers 930416.0800)

When I started this, I intended about 10 or 20 lines. It is now
three hours later. I have a lot of work that I intended to have
done in those 3 hours. Like Rick, I love the beauty of PCT, and
am addicted to it, but that, like any other addiction, can be
debilitating. I will probably fail, but I am going to try not
to write any more (except possible very short ones) until after
I return in June. Important as I think this is, its importance
to me is long-term, and the meetings I must prepare for are close
at hand.

Now to the apology to Bill, whom I honour too highly to offend
deliberately.

And what about your patronizing exasperation about my not paying
attention to the relationships you had laid out and that I had
not duly memorized?

Mea maxima culpa. I'm really sorry. I should know better than to
try to deal with 200 mail items at a time when I'm tired and frustrated.
But that's no excuse for rudeness. Only a possible reason.

The valid formulae are
Perceptual information per sample = log ((P+r)/r)
Disturbance information per sample = log ((D+r)/r)

This can't be the correct set of formulae. The implication is
that the perceptual information per sample is independent of the
(imaginary) sampling frequency. When you let the sampling
frequency go to infinity to approach the continuous case, the
information per sample must decrease accordingly, or else the
information rate will go to infinity. It will go to infinity-
squared, because r also approaches zero.

Since you can't get hold of Shannon's book, I'll xerox the appropriate
pages and send them to you. The formulae can be restated as log SNR for
P and for D. And it is true that if the SNR remains constant, and the
bandwidth goes to infinity, so does the information rate. But actually,
in any physical case, the information rate is not infinite, because the
faster the observation is taken, the less accurate it is. As the rate
goes to infinity, so does r (not to zero).

···

------------------
I don't understand why there is this rubber-banding between us. One moment
you write things with which I agree except for some apparently trivial
point, and the next we seem to be poles apart, working from wildly
different premises. I'll try to answer some of this, but there must
be a linguistic barrier somewhere. We have backgrounds similar enough
from the engineering side that we shouldn't have such difficulty talking
(except when good judgment is marred by frustration).

One sticking point seems to me to be the taking of appropriate viewpoints
on uncertainty. For example, you sometimes bring up the notion of a
sine-wave disturbance, and say it has zero uncertainty. But that is true
only of someone who knows that it is to be and to continue to be a sine
wave, and who knows the exact amplitude and phase of it. If any of those
factors are not known, then it has uncertainty. Uncertainty is a matter
of prediction. When a PIF is observing a sine wave, what the uncertainty
is depends on the PIF, not on the sine wave. What is the PIF capable of
observing? It does not anticipate anything, in the cases you usually
describe. It reports. So the information in the Sine wave for such a
PIF is exactly the same as for any other waveform with the same distribution
of intensity values. If the PIF were a phase-locked filter, it would be
different. The PIF would anticipate that the signal over the next little
while would continue the same sinusoid. A very narrow filter takes a
very long time to build up to the signal level, which it will do if it
is tuned to the right frequency. And it will take a very long time to
settle down when the signal disappears. It can accept information only
at a very low rate, commensurate with that of a nearly perfect sine wave;
it "knows" it is going to see a sine wave, and doesn't react to anything
else.

Once more unto the quantized breach, dear friends. Let's get it
straight. I do not now, and never have, belonged to the
quantized party of America. I don't plead the Fifth. I don't
have friends who are quantized. They all behave continuously,
like good citizens.

Then why do you keep talking about D/r? Why not just let r go to
the limit of zero (the continuous case) and the uncertainty to go
infinity (log(D/r), and use those continuous equations?

Hoo, Boy, do we have a misunderstanding! I let you go on about D/r
and used it because it seemed to help you to understand. But in the
continuous case, r does not go to zero at all. r is a measure of
resolution. It can be thought of as a noise power. D is a measure
of power in the D signal. (D+r)/r is more or less the continuous
equivalent of the discrete D/r.

It seems
to me that all the basic definitions require that D/r be a finite
whole number, and that the boundary between d and d+r be fixed,
with no values of d between d and d+r. To me that is quantization
-- the division of a continuous scale into fixed intervals.

That is quantization, but no basic definitions of uncertainty
or information require it.

That is to say, if the bandwidth of the disturbance is 1 hz, and
the bandwidth of the perceptual function is 1 Hz, the maximum
usable loop gain in the associated control system is zero. This
is simply not true. We will shortly have a new version of Simcon
that allows creating a random disturbance with a one, two or
three stage filter to permit varying the bandwidth and the
frequency rolloff. So you can create a disturbance with a
bandwidth of 1 Hz. Using an amplifier with a time constant of 1
sec (or whatever the right number would be) you can give the
control system an input bandwidth of 1 Hz.
...
You will find that the control system can be adjusted to keep the
effects of the disturbance on the perceptual signal very small,
with a loop gain very much larger than 0. If you don't want to do
it, I will do it for you.

Remember, when you do this, that you must not reduce the effective
bandwidth in the output section (as an integrator would do).
The argument applies to whatever portion of the disturbance is controlled,
not for the physical disturbance itself. It is why, when things are
moving too fast for us, we can improve matters by sitting back and
acting slowly. If you put a bandwidth limitation in the output, you
are defaulting on the control of the faster components of the disturbance.

I proposed this as a pretty strong test. You believe that you will get
good control under the following circumstances:

A simple ECS, with one sensory input, a constant zero reference signal,
a wideband very high-gain output that has an instantaneous effect on
the CEV. For "wideband" and "instantaneous" I am prepared to accept,
as a surrogate for infinity, a bandwidth an order of magnitude greater
than that of the PIF (Bp), and a delay an order of magnitude less than
1/2Bp. The PIF must be physically realizable (though simulated) have
an equivalent rectangular bandwidth that can be specified (most
computable filters have tabulated values, so that's not a problem),
and the disturbance must have a bandwidth of at least Bp (preferably
exactly Bp, but that could be hard to arrange).

You can do simulated-time samples at any rate you like greater than the
Nyquist rate of the output section.

If you find that you can obtain good control of the higher-frequency
components of the disturbance under these conditions, I will admit that
there is something wrong with my premises. If you don't, you will, I hope,
take my premises more seriously, even if I can't get the algebra right
all the time. (I've always known about this failing, which is probably
the reason why I tend to shy away from formula development. I'm much
happier with functions, and even worse with arithmetic than with algebra).

The bandwidth of the disturbance has absolutely nothing to do
with the maximum usable loop gain in a control system. There is
something wrong with your premises, which renders all the
manipulations that follow from them spurious.

Do you maintain this if the word "controllable" is inserted at the
beginning: "The controllable bandwidth of the disturbance...?" If so,
and if I am right, then I think IT will have begun to demonstrate its
value for PCT. If I'm wrong, then I've wasted a lot of net bandwidth
and my own energy, but I will have come to a point where I recognize
that a new insight is required. That's good for me, if not for the
rest of the community.

In passing, I'm curious why there is such rejection on the IT
side of concepts like variance and correlation, which involve
exactly the same kinds of derived measures:

There isn't a rejection. Correlation and variance are measures that
are useful in linear systems. So long as everything is nice and linear,
the analyses are pretty much interchangeable. But the information measures
are more general, and work in highly non-linear systems. For instance,
if the relation between x and y has even symmetry (e.g. x = y^2), and
y varies symmetrically around zero, there will be zero correlation between
x and y, but each will be derivable from the other, with the exception
of a 1-bit piece of information about whether y is positive or negative,
given x. There is almost perfect information about each in the other,
even with zero correlation between them.

As to "the uncertainty is that of the CEV given the perceptual
signal", that uncertainty has to be zero, because the CEV is
defined as the external correlate of the perceptual signal. You
are arguing from the premise that the CEV has an objective
existence independent of the control system, so that the
perceptual signal will represent it more or less accurately. Thus
the CEV could be in one state, while the perceptual signal
represents it as being in a somewhat different state. But that
contradicts the definition of a CEV.

If you define the CEV as the external correlate of the perceptual signal,
you are right. I haven't, so far as I know, ever defined it that way.
I define it as the external correlate of the PIF. The external correlate
of the perceptual signal is the state of the CEV, or its value, the way
I think of it. Anywhat, that's the sense in which you should read what
I have written. We must come to an agreement about this terminology.
A lot of the unnecessary discussion last month has been about mismatched
variables: adding output to disturbance and calling it perception, for
example. We do have to try to keep things in their proper domains.

From the Engineer's point of view, there is an inside and an environment

for an ECS. Inside, there is a perceptual signal. Outside, there is a
CEV that has a state. The state of the CEV as seen by the Engineer may
not always be reflected in the same value of the perceptual signal. There
is uncertainty about the value of the perceptual signal given the state
of the CEV, and about the state of the CEV given the value of the
perceptual signal.

The valid formulae are
Perceptual information per sample = log ((P+r)/r)
Disturbance information per sample = log ((D+r)/r)

This can't be the correct set of formulae. The implication is
that the perceptual information per sample is independent of the
(imaginary) sampling frequency.

That's right. It depends only on the RMS value of the signal and the
uncertainty of the measurement that the sample represents. The sampling
frequency is irrelevant, except that there is an informational limit,
in that samples taken closer together than the Nyquist limit are not
informationally independent.

Your formulae above do not take physical time into account, which
is why they lead to an absurd result as the sampling frequency is
raised.

The absurdity comes only from a misreading of the effect of changing
the sampling rate.

The formulae can't take physical time into account. They are IN physical
time. We aren't here talking about the samples taken at some arbitrary time
by a sensor looking at a waveform. We are talking about the real time
rate of decay of informational independence that is related to the
effective bandwidth of the signal. No way around it.

Without an explicit variable linking your discrete
calculations to physical time, you can't make the transition from
the discrete representation to the continuous (physical)
representation -- correctly.

Fourier figured out how. I'm only a follower here. I'm presenting
nothing new, and I don't think I'm misrepresenting the old and true.

The reason is that intrinsic delay, if you look from a
straightforward analogue signal-processing viewpoint.

I think you're confusing "disturbance bandwidth" with "the
highest frequency component of the disturbance." And note that a
pure delay doesn't change the bandwidth at all. Only an integral
lag (or some such) will cut off high frequencies. Integral lags,
by the way, introduce physical time.

Any physical filter introduces delay. That's intrinsic, and you can't
get around it. You can add extra delay, but you can't reduce the delay
inherent in the limited bandwidth of the filter. I'm sure you just
forgot it. The delay is of the order of 1/2B, if I remember correctly.

And I think you're still equating "disturbance" to "change in the
CEV." You've never understood why I make the disturbance, or the
disturbing influence, independent of the state of the CEV. The
reason is precisely to avoid the sort of confusion we have here.

The more different descriptions of disturbance we get, the more confusing
the discussion becomes. It has definitely been a problem. I brought
it up a few postings ago, and you said you had never changed your
definition. All the same I do get the impression you interchange
"disturbance" and "disturbing variable" quite a lot. But I don't
equate the disturbance with the change in the CEV, and don't remember
ever doing so, even in my earliest, most naive PCT days. I might well
equate it with the change that would occur in the CEV in the absence
of control, but I don't think even that is correct, given the non-linearities
of most real situations.

If the CEV is the position of a limb, the disturbance is not a
change in that position, but (for example) the magnitude of a
force applied to the limb.

This causes me a problem, for the reason of incompatible dimensionality
that I mentioned earlier. If the perceptual signal has as its external
correlate a position, how can a force be added to that position in the
computational algebra? Surely you have to include a transform based on
the mass and compliance of the object whose position is being perceived?
I would have thought that the "disturbance" would have had to be the
result of applying that force to the mass and compliance, a result that
would be at least in the dimension of position, although one would also
have the time dimension in it, so even that would not really be correct.

I do realize, of course, that the effects of the output in this situation
are forces, not positions.

And remember that the measure of the CEV is the
perceptual signal, by definition: the input function doesn't even
come into it.

This, I agree with. But it contradicts your definition of CEV, above,
where the perceptual signal defined the CEV, rather than measured it.
The input function comes into the definition of the CEV, not into the
measure of the CEV, though of course, given the value of the CEV and
the input function, our Engineer could determine the value of the
perceptual function.

I agree about the simulation test, but come on! If you don't take
responsibility for the correctness of your algebra -- on which
all of your arguments depend -- who will? You've already gone off
on a number of deductive tangents as a result of making errors.

Algebra can be checked by anyone, though it's much nicer to be able to
rely on oneself, and on someone else making a claim. I rely on your algebra
when you present it, because I assume that you are careful and get it
right. I try to be careful, but I KNOW I get signs shifted in copying
one line to another, and in this case I switched Bp/Bd into Bd/Bp. But
I don't know what deductive tangents I have gone on as a result of error
that were not fundamentally reasonable. I "see" the behaviour of the
system. The equations are a nasty necessity. Granted, I can "see" wrong,
so the equations ARE a necessity. But usually the deductive tangents
don't come directly from the equations. Think of the problem as analogous
to that of a colour-blind artist.

How can you convince anyone that your "principles" aren't equally
tainted by wrong deductions, when your defense of them depends
entirely on the mathematical manipulations you apply to your
premises, without any guarantee that your manipulations are free
of error?

If the manipulations were all there were to it, there is never a
guarantee. Bayes theorem says that the convincing is in the reader,
not in the writer. If a manipulation looks plausible and the result
doesn't disturb the previous belief, it will be accepted much more
readily than a complex one that refutes a belief. No guarantees. You
perceive what you perceive. If I leave you uncertain as to the validity
of an argument, by having been wrong in the past, you will need more
proof to change a belief on the basis of my argument in the future.
My problem, not yours. Your percepts, not mine.

Or is your faith in IT so unshakeable that no mere
mathematical proof or disproof (or simulation) could disturb it?

I think that if I understood the mathematical proof, and it seemed
correct, I would have no problem. If I understood that the simulation
satisfied appropriate conditions, I might doubt the theory, but I'd be
more inclined to look for problems in the additional constructs that
are necessary for the simulation. If a variety of simulations with
different types of boundary conditions all showed up a problem, I would
probably doubt my understanding of the theory, and would cease to use it.

As you said, the problem here isn't with the theory, but with its application.

As soon as we begin to treat the disturbance as a systematic
message instead of noise, we are out of the realm of
information theory.

NO.

YES. Why treat a continuous and regular phenomenon as if it is a
random variable? Why not use the simple and direct reasoning
appropriate to continuous variables when that is the kind of
phenomenon we are dealing with? Why not reserve the much more
complex and opinion-weighted and assumption-sensitive arguments
of uncertainty theory for situations in which the variables are
in fact unpredictable and irregular on the scale of interest?

We are both talking about continuous signals, using appropriate
methods. As to whether a variable is unpredictable and irregular,
that's ALWAYS a question of viewpoint. If the PIF is simply a
transformer, taking an input value and turning it into a neural
signal, all inputs are effectively irregular and variable to it.
But this is not true of the control system as a whole. It has its
own dynamic properties, which make it react differently to different
disturbances. A well-designed system may indeed be able to react
as if the variables were unpredictable and irregular, and thereby
minimize the fluctuations of the error signal. But another, built
for speed of response, may go into oscillation when hit with an
oscillating disturbance at the right frequency. For that control system,
the variables are more predictable and regular. It's a question of
viewpoint.

Nope. The uncertainty in the perceptual signal is determined by
the resolution of the perceptual apparatus.

So it's premise against premise. I like my premise better.
...
The resolution of the perceptual apparatus is infinite.
...
Thus the uncertainty in the perceptual signal is set by channel
noise, not by the resolution of the perceptual apparatus.

Definitional problem again. There's no premise here, but there may
be a language problem. For me, the perceptual apparatus of vision
includes the eye and the optic nerve, as a minimum; the perceptual
apparatus of hearing includes the ear, the hair cell, the auditory nerve.
All of these have resolution limits set in part by their physical
limitations and in part by the non-constancy of firing frequencies.
The channel noise is very much a part of that perceptual apparatus.

If the disturbance is a pure sine wave, it
contains no uncertainty at all, although one can treat the sine
wave using the manipulations appropriate to random variables and
get numbers out of the manipulations. The numbers, however, are
meaningless, because there is in fact no uncertainty.

I apologize for being boring, but we have another viewpoint problem:
"there is in fact no uncertainty" to whom? Not to the perceptual
apparatus that could equally well respond to a random disturbance. It
finds the sine wave exactly as uncertain as any other signal would have
been. To the experimenter who inserted the sine wave, there is no
uncertainty. And that is an uninteresting place for the uncertainty
to vanish.

The only uncertainty that exists in the disturbance is the
_unpredictable_ component, which can range from almost none to
the entire signal. The mathematics of uncertainty should be
applied ONLY to the unpredictable component. This has nothing to
do with whether the disturbance is opposed or unopposed.

Yes, to the component that is unpredictable at the point where the
uncertainty is of interest, typically at the perceptual signal line.
That, after all, is why we are ourselves control systems. If disturbances
and the effects of output were predictable, why control? We control
only the unpredictable component. The rest, we can predict and compensate
without control (cognitive planning).

The other [limit on control] is the perceptual inability to
determine what needs to be corrected, and in an effective
control system, this is the major limit, in my view.

It is not the major limit in our simulations. It is hardly a
consideration at all. And in the behavior that is matched by our
noise-free models, it can't be the major factor.

Are you sure? Look at what you write below:

If we lower the disturbance frequency enough, the person can keep
the cursor on the target with errors of only 1 or 2 pixels. So
the limit set by perceptual resolution (here enforced by the
screen resolution) is negligible with respect to the limits set
by dynamic considerations.

The limit is closely the limit imposed by perceptual restrictions,
whether or not that is affected by the screen resolution. Take the
subject 20 ft from the screen, and see whether they maintain a 1-pixel
tolerance. If that works, go 200 ft from the screen. I'll guarantee
you will find a distance beyond which the error will be a more or less
constant angle subtended at the eye.

There's no argument that there can be limits other than perceptual, which
overwhelm the perceptual limits. Inadequate musculature, bad dynamics,
etc. etc., all can make matters worse. But even in a well-functioning
control system, one CANNOT control what one can't perceive. And I don't
really believe that you are saying that one can. But to refute what I
have been trying to get across, that is just what you must do.

It seems to me that the reasoning involved in information theory
depends to an inordinate degree on taking just the right attitude
toward it, making just the right interpretation (on which even
you and Allen don't seem to agree and which others versed in IT
like Cliff Joslyn also see differently).

I know that there is a lot of misunderstanding of IT. Much of it hangs
on the telephone metaphor, as I commented to Mary the other day. The
idea that information is something that passes through a channel from
one end to the other. That's a use of information theory, but it is
a bad metaphor. PCT is based on the idea that you have to take an
internal viewpoint--the perception you control is YOUR perception, not
anyone else's, or that of an abstract model. I have been trying to hold
you to the same standard with information theory. There are disagrements
between you and non-PCT psychologists because they (unknowingly) use
information that is not available to the person studied. There are
disgreements between me and (some) information theorists for exactly
the same reason.

I don't think there is much disagreement between Allan and me, except in
expository style. Sometimes even I don't understand him, but since we
are located in the same corridor, I can get a verbal explanation. He
is expounding at a different level of abstraction from me, but what he
is actually expounding in his NP-hard discussion is what we call "scientific
method." It is what you DO when you make models of control systems.
The results are Occam-good. What I have been trying to do is different.
I make assumptions about what a particular system can or cannot know,
and infer from that how it could or could not behave.

And I am definitely not impressed with the
manipulations I have seen so far, which are loaded with errors
and erroneous predictions.

Loaded with errors is a bit strong. I'm vulnerable to the charge of
poor algebraic manipulation, but you have not shown one erroneous
prediction (yet), except for those implicit in the results of the
erroneous manipulations.

So far your principles have not done
well by you in terms of leading you to verifiable statements
about either real systems or simulations.

Perhaps by the time I return, the verifiable statement about which
we disagree will have been refuted or confirmed. If you remember, I
made a previous prediction, and suggested an experiment which Rick
(not you) accepted as a very good strong test. He was sure (and I
seem to remember you were, too) that we would be surprised by the
result, and would have to rethink our position. When it came out as
I anticipated it would, it suddenly became a nothing test, all shot
with false assumptions. But those false assumptions were not mine
(or Allan's, who constructed the actual test). We stated the assumptions
by quoting directly from you and Rick, and asked whether those assumptions
were still held. We were told that they were, until AFTER the test, when
they were no longer valid. So I think your statement here is also a
little strong.

Check out the Gain-Bandwidth expression, and see how it works.

My wife just called to find out why I am so late for supper! Another
perceptual signal not properly aligned with its reference.

Sorry for the length. As Victor Hugo said, I hadn't the time to make
it shorter.

Martin