[Martin Taylor 970427 17:05]
Bill Powers (970426.2210 MST)]
We reallyt must live on different planets. I can't make head nor tail
of half of your posting. But we forge ahead, regardless, through the murk
(Incidentally, I won't be able to do much more of this before leaving
for a month at the end of next week. I've suddenly been told I have
to prepare a presentation, and I haven't even thought about the material
yet.)
Martin Taylor 970426 21:45--
I deduce that you think that the informational analysis in
some way presupposes that the control system separates out _two_ input
signals, one based on the disturbance signal and one on the output signal.
Is that correct?
Yes.
If you are supposing this, then unsuppose it. The analysis needs no more
wires, variables, or signal lines than are in the normal everyday diagram,
and "p" is still a scalar variable that is a function of qi.
But you assume that there is available some measure of the disturbance
signal other than qi itself. For the analyst this is true -- the analyst
knows qd and Fd, or can at least assume an equivalent qd and Fd. But for the
control system it is not true: the control system knows only qi.
That's correct. When you are analyzing a control system, you are an analyst,
aren't you? You want to know, for example, how much the action of the
control system has reduced the effect the disturbing variable would otherwise
have had on the input variable. You may want to know other things, such
as the spectrum of the error signal. The control system doesn't "know"
about that, does it? Nor does it "know" qi. The analyst knows qi, and
the analyst knows p, the perceptual signal that is a function of qi. The
analyst knows the function, the control system doesn't. The control system
only behaves the way it behaves, having the functions it has, and being
provided with the inputs it is provided with--the reference signal and
the disturbance signal.
I did demonstrate once that it is _possible_
to reconstruct the disturbance given no variables except the qi waveform,
using the fixed function that happened to be Fe(Fo()). That demonstration
only showed that the perceptual signal does carry a lot of information
about the disturbance, not that this information is segregated in the
action of the control system.
But you have also said that the better the control, the less information the
perceptual signal carries about the disturbance signal, with a limiting
amount of 0 for perfect control. This is why I keep saying that you have not
yet really solved this problem: what you are saying in words, and mostly
from personal intuition, needs to be formally derived in mathematics. If the
amount of information in p about the disturbance signal is a function of the
quality of control, you need to prove that and quantify the statement.
What is your measure of the quality of control? There are quite a few, I
believe, but they all come down to the reduction in influence the disturbing
variable has on the perceptual signal. Informationally, that concept is
equivalent to reduction of the mutual information between the perceptual
signal and the disturbance signal.
And by the way, since we have agreed that the default meaning of disturbance
is qd, I do wish you would use your own term and consistently say
disturbance _signal_ instead of just disturbance. The two terms mean
different things.
I try, but when the meaning is obvious, but it's hard to keep using the long
form--rather like forgetting to say "daily newspaper" and using "paper"
instead, when "paper" is already used for something else. But I'll
continue to try.
Given all the functions and variables in the loop (qi, Fi, comparator, Fo,
and Fe), plus the reference signal, it is possible to compute qo. Once you
know qo, you can deduce the part of qi that is due to qo, the remainder
presumably being due to some disturbance signal. But there is no independent
way to determine the disturbance signal, unless you know qd and Fd.
In {Rick's] terms, what is needed is f() and g(d), not f(u) and g(d).
Yes, because u (qo) can be calculated from qi, Fi, comparator, r, Fo, and Fe.
That's right. That's why it works. The point of the demo was that there
was no VARIABLE RELATED TO g(d) other than the perceptual signal, and
nevertheless the fluctuations in g(d) could be reconsituted pretty accurately.
The quality of the reconstitution is the same as the quality of control.
This proves that the fluctuations in the perceptual signal convey all
the information from the disturbance signal that is used in control. The
fact that to make the reconstruction requires various functions and a variable
that is independent of g(d) in no way alters that proof. They don't have
any fluctuations related to g(d). Only the perceptual signal does.
But it's rather irrelevant that an analyst _can_
do this kind of partitioning, except to show that the information about
the disturbance signal in the perceptual signal is not lost--at least the
part that has actually contributed to control is not. The rest shows up
as the correlation between the disturbance signal and the perceptual >signal.
You are speaking from intuition here; I repeat, your statements won't mean
anything until you have verified them by a mathematical derivation.
The last sentence might be constured as "intuition". But even _you_ can't
assert that a high correlation between the perceptual and disturbance signals
doesn't show failure of control. Or can you?...I sometimes wonder just what
strange things you _can_ say about correlation.
------------------
On correlation and bits
I presented a table that would be
part of it, the part you have in the past been interested in, and that
represents the degree to which control is imperfect.
But this table is good only for open-loop calculations. I suspect that
closed-loop calculations will provide some surprises.
If you look at it, you'll see that it has nothing to do with open or
closed loop calculations. It has only to do with the values of data.
The value of the bit rate calculated from the correlation is correct
for bivariate Gaussian distributions, but for other distributions
that calculation understimates the bit rate (I'm pretty sure but not
certain).
But if y = f(x,y), the correlation you calculate between y and f(x,y) is not
the correlation between x and y.
Agreed. The correlation between y and f(x,y) is 1.0.
... open and closed loop are merely mechanisms whereby correlations come
to be as they are. The correlations themselves map onto information
transfer, or rather, they provide lower bounds on the information
transfer between the correlated variables.
That's what you say now, without having actually derived the correct
equations.
Richard Kennaway did that. The correct equations have to do with what
value of x occured along with what value of y on this occasion, that
occasion, and the other occasion(s). There's no issue of how the values
were produced, just what the values turned out to be. You can calculate
the correlation between the number of hairs on a dog and the number of
sunspots on the visible part of the sun when it was born, if you want.
There's a difference between calculating the correlation in a data set
consisting of pairs of numbers, and calculating it based on a theoretical
model.
There's certainly a difference between predicting a certain value of
correlation and finding it in the data. That I'll grant. But what does
this have to do with anything?
If you see that y = 10*x, you would expect a correlation of 1.0
between x and y. But if, unknown to you, x = y^3, the actual theoretical
correlation, and the correlation you will measure, will be zero, because the
solution is y = 100, a constant.
??? This makes absolutely no sense to me. It's a word jumble.
Calculations of correlation are done to determine whether there is
_linear_ independence between the variables. You are presented with a set
of data pairs ({x1,y1}, {x2,y2},....{xn,yn}) and you calculate a
correlation.
Or you can calculate a mutual information measure. Or you can assume that
the pairs come from a bivariate Gaussian distribution and compute a
lower-bound bit rate from the correlation. What assumptions were you
thinking of?
I was thinking of theoretical relationships, such as in your example of
Y=|X|, as in the following paragraph. If X depends on Y by some other
pathway, so there is a closed loop, the theoretical correlation is no longer
the same.
If Y = |X|, then Y = |X|. It doesn't matter what else is true, it is true
that no matter what value you measure for X, the value of Y is its absolute
magnitude. You can't even _get_ a correlation until you have measured
several different values of both. We know from the specification that
X is symmetrically distributed around zero. That specification ensures
that there is zero correlation between X and Y, even though knowing X
enables you to know Y.
To see this latter point, imagine a relationship Y = |X|, with X having
a zero-centred normal distribution. Knowledge of X provides perfect
knowledge of Y (in info-theory terms, the mutual information of X and
Y equals the uncertainty of X) but the correlation between X and Y is
zero.
But what if Y = |X| while X = k*(|Y0 - |Y}||)?
There's some problem with your notation here, but it doesn't matter,
provided that whatever you intended is logically consistent with the
specifications, namely Y = |X| and X is from a zero-centred distribution.
I was trying to introduce a second relationship that creates a closed loop.
If X depends on Y at the same time that Y = |X|, your statements about
"perfect knowledge" and so forth do not hold.
Of course they hold. Each time you get a value of X, you find that the value
of Y is X if the measured X turns out to be positive, and is -X if the
measured X turns out to be negative. That's perfect information about Y
given X.
Let's try and see if this is consistent with the
specification of the situation, which is: Y = |X| and X is symmetrically
distributed about zero.
Substituting for Y in your expression, we have X = k*(|Y0|-|X|).
Collecting, we get X - k*|X| = k*|Y0|. Assuming Y0 and k to be fixed,
this has only one solution for X, doesn't it? And that's inconsistent
with the specification that X is symmetrically distributed about zero
(in the original, it was more precise, a zero-centred Gaussian), so this
formula is invalid.
The formula is valid if, in fact, there is a feedback path from Y to X;
Well, there can't be, can there, if to have one is inconsistent with the
specifications. But I wouldn't be surprised to find you _could_ construct
a feedback path that would allow the demonstration data to exist. I leave
that as an exercise for the reader. It doesn't interest me.
in
that case it must be your assumption that X is symmetrically distributed
around zero that is inconsistent with the possible behavior of the system,
or else we are seeing a degenerate case (the symmetrical distribution has a
standard deviation of zero).
You are really strange. I provide an example of data in which a zero
correlation goes along with high mutual information, to show that the
correlation measure is only a lower bound on I(X|Y). You construct
a circuit in which the demonstration data won't occur, and use it to
argue that therefore the arithmetic is wrong. I don't understand this
mode of argument, unless it is a rhetorical trick to ensure that the
point is lost in the fog. If that's what it is, I understand it and don't
like it.
If you prefer, I here is a list of data in which Y = |X| and ask you to
compute the correlation. You can construct whatever circuit you like that
might have generated those data, but if your circuit wouldn't have generated
them, it will be an irrelevant circuit.
X = 1, 3, -3, -2, 1, 4, -1, -1, 2, -4
Y = 1, 3, 3, 2, 1, 4, 1, 1, 2, 4
Read Section 5 of Kennaway's paper, if it would help.
Here's another point to mull over. If X provides B bits of information about
Y, Y provides B bits about X. It matters not a whit whether X influences Y,
X and Y are both influenced by Z, Y influences X, or none of these are true.
Information, like correlation, carries no implication of causality.
Martin