Importance of disturbance discussion to PCT

[From Bill Powers (970422.0939 MST)]

Why is it so important to settle this discussion of "information about the
disturbance in perception"? Because this is the critical point in
distiguishing the PCT view of behavior from all theories in which the
environment ultimately causes behavior.

The "behavioral illusion" is the basic phenomenon that is treated
differently in PCT. Something happens in the environment -- some physical
variable in the environment changes its state. Closely following that
change, we see a change in a measure of the behavior of an organism. The
appearance is that the physical variable had some effect on the organism,
causing it to produce the change in behavior. This phenomenon has been
interpreted in the past as evidence that environmental stimuli act on the
sensory inputs to an organism, causing a series of neural events inside it
that end with muscle tensions and the movements they cause. Thus it seems
that something about the stimulus itself, the external physical event, is
the cause of the behavior.

PCT leads us to suspect that this is not the correct view. PCT proposes that
most often, perhaps even in all cases, there is a controlled variable that
is affected in one direction by the environmental physical variable, and in
the opposite direction by the change in behavior. Furthermore, the change in
behavior is caused _entirely by the change in the controlled variable_, and
is completely independent of the nature of whatever it was that caused that
change. It is "independent" in that any environmental change that can have
the same perturbing effect on the controlled variable will be met by the
same change in behavior. Different physical variables can cause the same
change in behavior without having anything in common with each other --
except their common effect on the controlled variable.

Furthermore, this "common effect" is dependent on the perceptual
organization of the control system, not on the actual physical situation
outside it. If the perceptual organization changes, two physical variables
which formerly led to the same behavior may now lead to different behaviors;
or only one of them, or neither, may now lead to changes in behavior. The
controlled variable is not defined in terms of the physics of the
environment, but in terms of the perceptual organization of the control
system. Hence, the "common effect" is not physically determined by anything
outside the behaving system.

From this, it follows that there is no way to reason backward from the

behavior of an organism to the environmental events that gave rise to it. A
change in behavior can be traced back to a change in the controlled
variable, but that is where the trail stops; there is no way to associate
particular behaviors with particular causal events in the environment. The
reason is that there is an infinity of _different_ events in the environment
that can all lead to the same behavior, and there is no basis for choosing
any one of them as "the" cause. Using only information available to the
control system, we may be able to say that there has been _some_
disturbance, some change in an external physical variable, but we cannot say
in general what it was.

Or more important (since in specific cases, as external observers, we know
everything about the local environment), the _control system_ does not
contain any information specific to the environmental cause. The point is
that it does not need such information. It senses and acts upon the
controlled variable, and to do this it does not need to know why the
controlled variable has changed.

I should note that even in the S-R view, there is no justification for
saying that something about the physical stimulus or stimulus event causes
the behavior. The same argument applies: only the proximal effect on the
sensory nerves can cause anything in the organism, and that proximal effect
could arise from many different distal variables. I think this fact is often
overlooked by experimenters who speak of "The effect of X on behavior."

Best,

Bill P.

[Martin Taylor 970422 14:30]

Bill Powers (970422.0939 MST)]

Why is it so important to settle this discussion of "information about the
disturbance in perception"? Because this is the critical point in
distiguishing the PCT view of behavior from all theories in which the
environment ultimately causes behavior.

Huh? Everything you say following this is the common ground from which
we both (all?) work. How can it relate to the issue of "information
about the disturbance (Fd(qd)) on perception" if it is agreed to be the
foundation for the arguments of both "sides" of the discussion?

The key statement is:

Or more important [...], the _control system_ does not
contain any information specific to the environmental cause. The point is
that it does not need such information. It senses and acts upon the
controlled variable, and to do this it does not need to know why the
controlled variable has changed.

Right. Roger and Over.

We all know this. Always have. Now can we get on with the issue of the
information about the disturbance (of the input quantity) in perception--
or rather, the issue of the informational analysis of the behaviour of
the control loop?

Or should we drop it again, on the grounds that the trail will soon be
undetectable again among all the red herring trails, after having seemed
so promising only a week ago?

Martin

[From Bill Powers (970422.2011 MST)]

Martin Taylor 970422 14:30]--

Why is it so important to settle this discussion of "information about >>

the disturbance in perception"? Because this is the critical point in

distinguishing the PCT view of behavior from all theories in which the
environment ultimately causes behavior.

Huh? Everything you say following this is the common ground from which
we both (all?) work. How can it relate to the issue of "information
about the disturbance (Fd(qd)) on perception" if it is agreed to be the
foundation for the arguments of both "sides" of the discussion?

Yes, it is common ground and that is why I wrote it, and where I want to
leave it. You don't need my permission to pursue your interest in
informational analyses of control systems; I don't need yours to pursue my
interest in PCT. If you want to persist in calling Fd(qd) "the disturbance,"
when that is NEVER what it has meant in PCT before you came on the scene,
there is nothing to stop you. I will continue using it the way I always
have, and in that meaning it is perfectly correct to say that there is no
information in the perceptual signal about the disturbance (except, as you
have pointed out, in special circumstances and then only in the eyes of the
analyst). If you object to my saying that, it is only because you mean
something different by "the disturbance."

We all know this. Always have. Now can we get on with the issue of the
information about the disturbance (of the input quantity) in perception--
or rather, the issue of the informational analysis of the behaviour of
the control loop?

For me, it's not an issue. It's you who are interested in pursuing the
informational analysis of the behavior of the control loop; I am only a
mildly interested bystander with regard to that subject. There is nothing I
can contribute to it. I wish you well in your investigations, but I will not
try to participate in them.

I ask you most earnestly, however, to try to find a different term for
Fd(qd), so when people read about your studies of the information in the
system that comes from Fd(qd), they will not think your statements are
contradictory to mine as I describe the relationship of disturbances to
variables in the control loop, using the term "disturbance" to mean qd -- as
I always have done.

Best,

Bill P.

[Martin Taylor 970423 10:00]

Bill Powers (970422.2011 MST) and others

For me, it's not an issue. It's you who are interested in pursuing the
informational analysis of the behavior of the control loop; I am only a
mildly interested bystander with regard to that subject. There is nothing I
can contribute to it. I wish you well in your investigations, but I will not
try to participate in them.

Fine. I'd prefer your constructive participation, but if you don't want to,
it's not a problem. Perhaps if I manage to develop some testable prediction
you might not otherwise have made, you may want to re-engage. If not, we
are where we have been stuck for a long time.

I ask you most earnestly, however, to try to find a different term for
Fd(qd), so when people read about your studies of the information in the
system that comes from Fd(qd), they will not think your statements are
contradictory to mine as I describe the relationship of disturbances to
variables in the control loop, using the term "disturbance" to mean qd -- as
I always have done.

I know. We've talked about the need for this term off and on for some years
(five, isn't it?), and whenever we seem to have agreed on one, we use it
for a while and then somehow its meaning becomes problematic again. last
week you introduced "perturbation." For a long time I've tried to be
consistent in using "disturbing influence" (Fd(qd)) as opposed to "disturbing
variable" (qd), but sometimes I slip and use the short form where (I think)
no confusion is possible.

A control loop has two inputs from and two outputs to the outer world.
Of these four, three have names that seem to cause little or no confusion:
"perceptual signal" (to higher levels), "reference signal" (from higher
levels), "output signal" (to the environment) and ?? (from the environment).
Could we perhaps call "??" "disturbance _signal_" to corresponds with the
other signals? Would that help avoid the confusion and unnecessary
argumentation?

Certainly when we are doing any kind of analysis of the loop, algebraic,
Laplace, Fourier, informational or whatever, it is Fd(qd) that enters, never
qd. Just as it is r, the reference signal value, not the various signals
from higher levels that have to be combined to form r. So I propose to
use "disturbance signal" and "disturbance signal value" (disturbance value
for short) for the environmental external input to the control loop,
unless that upsets you.

···

-------------------

I don't accept your use of "d" in your diagram because it is confusing to
people trying to understand the standard PCT diagram.

Why don't you accept it when I write it but are quite happy when Rick
uses the same symbol for the same signal?

--------------

(earlier, to Hans Blom) In fact you're doing the same thing
Martin Taylor did: confusing the _cause_ of a perturbation with the
perturbation itself.

I wish you would quit this repetition. It's must be five years since its
falsity was first pointed out to you, and apparently accepted by you. It's
almost four years since we clarified it again (as I thought) in face to face
conversation, and you repeat it over and over. Please, please, please
stop it. I've NEVER, NEVER, NEVER, had that stupid, unbelieveably stupid
idea. I've never said it, never done an analysis using it, never thought it.

You, for reasons obscure to me, initially asserted that's what I meant when
I said it wasn't, and for years, by repeating it, you have continued to tell
people that I am an idiot. Don't. It's insulting and unnecessary. And,
for reasons you might intuit, annoying. And it gets in the way of serious
discussion.

The reason I never thought it, as I pointed out only a few days ago for
the Nth time, was that when I first learned of PCT I considered only the
signals and parameters of the control loop itself. It never occurred to me
to consider the external causes of these signals (except, perhaps for the
causes of the reference signal, when analysing the behaviour of a hierarchy).
When you brought up the notion of external causes being "the disturbance",
that was a surprise, but not as much of a surprise as your attribution to
me of the notion that the control system could determine and take advantage
of the cause. It was never true, and doesn't become more true by constant
repetition.

Please stop.

---------------------------
+ Rick Marken (970421.0900)

+qi = f(qd+qo). So qi might have perfect information about the sum,
+qd+qo, but it has no information about the contribution of qd to
+that sum.

Bill Powers (970422.1622
I'll agree that Rick's point is different from mine, but it is not wrong. If
the perceptual signal is o+d, there is no way for the control system to
partition it into the part due to o and the part due to d;

Mutual information is not the same as partitioning.

On whether any information can be gained about a component from measurement
of a sum, have you actually read Richard Kennaway's recent postings, which
Rick praised so highly? Kennaway shows precisely how much information
can be gained about a component from the sum. Using his formulae, I
showed yesterday the rates for two signals having different correlations.

I hoped you would take these numbers and consider them in the case where
one signal is p and the other d (disturbance signal value). Instead, you
choose to hold the self-contradictory position that Kennaway's paper and
postings are fine, but his conclusions are not.

If two measures have a correlation c, then each measure of one of them
conveys -0.5*log2(1-c^2) bits about the other. If X = Y + Z, the correlation
between X and Y depends on the variances of Y and Z in the way Richard
described (if Y and Z are correlated, as they are in the case of o and d,
then there's a cross product term that comes in as well, but that doesn't
affect the relation between correlation and mutual information, at least
for gaussian variables).

-------------
In the same vein:

I think Richard Kennaway dealt with this interpretation very well.

Yes. He did accept what I said, after all;-)

You're
saying that the experimenter measured the X's, and just happened to combine
them so that Y was predicted perfectly. If the law (known only to God) had
been Y = aX1 + bX2 + cX3 + dX4, the experimenter would have had to guess at
a, b, c, and c in order to get the perfect prediction. His guess that Y =
sum of X's would not predict very well at all.

Bruce's example involved no guess about a, b, c, and d. A multiple
regression analysis discovers an estimate of those values, with a precision
that increases with the number of observations of the X's and Y. The rest
of your comment on this matter is therefore invalid.

Your scenario assumes (1) that the experimenter measures the _true_ values
of the variables, with no noise, and (2) that the experimenter combines them
as God combines them to produce Y. If these conditions hold, then it makes
no difference that the individual variables happen to vary with a normal
distribution, or that the experimenter is unable to predict their values in
advance. If you truly treat these variables as random variables, then they
have a mean and a distribution. If the true value of Y is the sum of the
true values of X, and the X's are simply normally distributed variables with
a mean of 0, then Y = 0 is the correct prediction.

What on earth are you talking about? It seems to have no relevance to the
question of what the distribution of Y is after you have measured the X's.
My "scenario," as you call it, assumes that the experimenter measures the
X values in exactly the same way as when evaluating the correlations in the
first place. Who knows the "true" values? And what does it have to do with
the argument?

Of course, if Y is normally distributed with mean zero, then before you
gain any further information about it, your best bet is that any particular
measure will have a mean of zero. After you have gained some information
about it, by measuring one or more of the Xs, your best bet will be
some other value.

Once you've done your regression analyses, and found that to your best
estimate Y = 1.5X1 + 2.3X2 + (-0.8)X3 + 0.15X4 + (random with sd 0.5), then
after you have measured X1 to be 2.0 on a specific occasion, your best
estimate of Y is 3.0, and it will have a standard error of estimate a
bit less than that of your original zero estimate. Assuming that the
variance of each X is unity, I think that the variance around your estimate
of Y will drop from (1.5^2 + 2.3^2 + 0.8^2 + 0.15^2 + 0.5^2) to
(2.3^2 + 0.8^2 + 0.15^2 + 0.5^2). Could be wrong there, but I think not.
After you measure all the Xs for this particular Y, your standard deviation
will be just the 0.5 which you earlier obtained as a consequence of your
errors in measuring the Xs and the Y while you were doing your regression
analyses. Your unknown errors in measuring the Xs this time are subsumed
in this.

-------------

For example: X1 = 3m +- 1mm, X2 = 2m +- 1mm, X3 = 4m +- 1mm, X4 = 1m +- 1mm

Y = 10 m +- 2mm. The X's are known to within about 1 part in 2.5 thousand
on average. Y is known to one part in 5 thousand. Not bad, and a little
better than the Xs individually. I don't know what the correlation between
Y and the sum of Xs would be in this system (measured with a precision
easily achieved with a ruler), but I'll bet it's over .995. If the Xs
are known to be positive, as is often the case, then Y will always be
known to better relative precision than the Xs individually.

Oh, Martin, you are such a slippery character. Your assumption that the X's
are positive is, of course, necessary in order to make your conclusion true,
but it is in no way required for or implied by Bruce's illustration. Let me
give you a different numerical example, one not selected to make your
conclusion be true:

X1 = 3m +- 1mm
X2 = -2m +- 1mm
X3 = -4m +- 1 mm
X4 = 3m +- 1mm

Sum: Y = -0m +- 2 mm

Now Y is known with a relative precision less than the average relative
precision of the Xs -- considerably less!

On "slippery," perhaps pots should not apply labels to kettles.

I have no intention to be slippery, or to present anything misleading. My
objective, though it may sometimes seem otherwise, is clarity. In this
case, I used the special case of positive measures to avoid unnecessary
complication while making a general point. Real measures typically have
a natural zero or are translationally symmetric; in the latter case the
zero is arbitrary. If a measure has a natural zero, the error of measurement
is usually something like proportional to the magnitude of the measurement.
If it doesn't have a natural zero, then the relation of the measurement
uncertainty to the magnitude of the measure is not legitimate. I didn't
see any need to introduce this kind of complication into the example.

If you want to do it right, you can work from the uncertainties of the
X values and the Y value before and after the measures are made. If the
standard deviations of your estimates of the X values before the measure are,
say 1 unit, then that of the Y value before the measure is 2 units. You
measure each of the X's with a precision of 0.1 unit, and then your prediction
of Y will have a standard deviation of 0.2 units, the same ratio. If the
expected mean sum of the Xs is different from zero, then it has to come
into play, always in the sense of reducing the relative error in Y as compared
to the Xs, as in my simplified example. Not so slippery?

Too long. Written in a state of too much annoyance:-(

But anyway, I continue to do what I can to try to understand PCT better,
and to help other people to do so.

Martin

[Hans Blom, 970424d]

(Bill Powers (970422.0939 MST))

The "behavioral illusion" is the basic phenomenon that is treated
differently in PCT. Something happens in the environment -- some
physical variable in the environment changes its state. Closely
following that change, we see a change in a measure of the behavior
of an organism. The appearance is that the physical variable had
some effect on the organism, causing it to produce the change in
behavior. This phenomenon has been interpreted in the past as
evidence that environmental stimuli act on the sensory inputs to an
organism, causing a series of neural events inside it that end with
muscle tensions and the movements they cause. Thus it seems that
something about the stimulus itself, the external physical event, is
the cause of the behavior.

Much of the problem that you consider is due to a confusion about or
different interpretation of the word "behavior". Nothing new, of
course. The standard interpretation of "behavior" is in terms of
observables: the organism's or controller's _actions_, i.e their
effect on the world. It is only those that we can observe, normally,
in organisms. We cannot look into somebody's head and have a direct
view of their goals/references. That organisms have goals is known
only by inference -- through a formal or informal application of The
Test. And you know how cumbersome and endless a procedure that is,
before we are truly satisfied that we have discovered a goal. So it
is not so strange that most people prefer to equate "behavior" with
observables (perceivable actions).

You want to see behavior as what follows from (internal) (changes of)
reference levels. So you put the less certain and more abstract thing
(references) first and the more certain and more concrete things
(observable actions) second. Some people would say that this is a
logical error: explaining more certain things by less certain ones.
But I agree with you that this point of view has its uses.

But although I sympathize with your point of view, it has proven to
be extremely hazardous to redefine a term which has a certain well-
established connotation. Too often people seem to interpret the ":"
in B:CP as an "=", you yourself not excluded. It remains my personal
preference to use the word "behavior" for _observables_. In HPCT that
would be the collective outputs of the lowest hierarchical level.

If we equate "behavior" with externally observable actions, we can
clearly see that in a control system "behavior" has _two_ "causes".
Take the theodolite controller: its output is determined by (1) its
internal reference level(s) _and_ (2) the external "disturbance".
When the disturbance suddenly goes to +10, the compensatory action/
output goes to -10. This is where stimulus-response theory applies --
_because a controller is at work_. But changes of the reference level
translate into output changes as well. And stimulus-response theory
therefore posits "internal" stimuli (e.g. hunger or thirst) that the
observables can be attributed to. Not unreasonable: they are the
higher level references of an HPCT hierarchy.

I think that as long as we keep in mind that externally visible
"behavior" has two causes, a fairly straightforward "translation"
between SR-theory and control theory is possible -- although control
theory gives additional insight. It is therefore, maybe, that PCT
impresses many people as "nothing but" something very much like a
top-down versus bottom-up coordinate transformation: useful, but
"nothing new". Alas: they miss the additional insight...

Greetings,

Hans

[From Bill Powers (970424.0751 MST)]

Martin Taylor 970423 10:00]--

So I propose to
use "disturbance signal" and "disturbance signal value" (disturbance value
for short) for the environmental external input to the control loop,
unless that upsets you.

It's OK with me, although I could quibble. I'm content if you will never say
"disturbance" without meaning qd.

-------------------

I don't accept your use of "d" in your diagram because it is confusing to
people trying to understand the standard PCT diagram.

Why don't you accept it when I write it but are quite happy when Rick
uses the same symbol for the same signal?

Because I know that the default meaning when Rick refers to the disturbance
is qd, whatever symbol he uses for it. Your default meaning is what you now
propose to call the disturbance signal. When Rick says "o + d" I
automatically translate, as he does, to "Fd(qd) + Fe(qo)", with the
understanding that the two functions are, for this simplified case only,
unity multipliers.

--------------

(earlier, to Hans Blom) In fact you're doing the same thing
Martin Taylor did: confusing the _cause_ of a perturbation with the
perturbation itself.

I wish you would quit this repetition.

If we are now in agreement about the default meaning of "the disturbance," I
will happily desist.

···

------------------------------

Bill Powers (970422.1622

I'll agree that Rick's point is different from mine, but it is not wrong.
If the perceptual signal is o+d, there is no way for the control system to
partition it into the part due to o and the part due to d;

Mutual information is not the same as partitioning.

But partitioning by some means is necessary if explicit use is to be made of
the information -- if the part of the perceptual signal magnitude due to the
disturbance is to be treated any differently from the part due to the
system's own action. The comparator compares the _whole_ perceptual signal
to the reference signal, not just the part due to the disturbance.

I do understand about superposition as a principle of linear analysis. In
fact, you might make some progress on this problem if you were to
conceptually split the basic control loop into two independent loops, one
dealing only with the effects of the disturbance signal and always having a
reference signal of zero, and the other experiencing no disturbance and
receiving the actual reference signal. The two systems, operating in
parallel, would have exactly the same form of output function and
environmental feedback function. The real state of the controlled variable
(also duplicated) is the sum of the states of the variables in the two
parallel systems, and the real output is the sum of the outputs. This would
mean that you can strip out just the closed-loop effects involving the
disturbance signal, without the complication of a variable reference signal.
That might make the closed-loop informational analysis a lot less messy. In
fact, with a reference signal of zero, you could combine the perceptual
function, comparator, output function, and environmental feedback function
into a single equivalent function and reduce the problem to a very simple
diagram. This might be the key to solving the problem of information in a
closed-loop system, and that would be interesting to see.
---------------------------

On whether any information can be gained about a component from measurement
of a sum, have you actually read Richard Kennaway's recent postings, which
Rick praised so highly? Kennaway shows precisely how much information
can be gained about a component from the sum. Using his formulae, I
showed yesterday the rates for two signals having different correlations.

I hoped you would take these numbers and consider them in the case where
one signal is p and the other d (disturbance signal value). Instead, you
choose to hold the self-contradictory position that Kennaway's paper and
postings are fine, but his conclusions are not.

No, I simply accepted his computation of the number of bits, and understood
how a one-bit prediction amounts to the prediction of only the sign of a
relationship. I haven't seen your computation of the number of bits
representing the disturbance signal as they relate to the information in the
perceptual signal. This case is complicated seriously by the fact that the
flow of information through the input function is considerably reduced by
the negative feedback effects of the output -- which is only to repeat,
without much understanding, what you have said several times. If you come up
with a valid computation, I will probably accept that, too. See the above.

If two measures have a correlation c, then each measure of one of them
conveys -0.5*log2(1-c^2) bits about the other. If X = Y + Z, the correlation
between X and Y depends on the variances of Y and Z in the way Richard
described (if Y and Z are correlated, as they are in the case of o and d,
then there's a cross product term that comes in as well, but that doesn't
affect the relation between correlation and mutual information, at least
for gaussian variables).

That's for the open-loop case, which is what Richard was discussing. The
closed-loop case, for all I know and for all you have demonstrated, may be
completely different. So get busy.

-------------
In the same vein:

I think Richard Kennaway dealt with this interpretation very well.

Yes. He did accept what I said, after all;-)

You're
saying that the experimenter measured the X's, and just happened to combine
them so that Y was predicted perfectly. If the law (known only to God) had
been Y = aX1 + bX2 + cX3 + dX4, the experimenter would have had to guess at
a, b, c, and c in order to get the perfect prediction. His guess that Y =
sum of X's would not predict very well at all.

Bruce's example involved no guess about a, b, c, and d. A multiple
regression analysis discovers an estimate of those values, with a precision
that increases with the number of observations of the X's and Y. The rest
of your comment on this matter is therefore invalid.

No, Bruce has now acknowledged that there is no uncertainty in the measures
of the X's or of Y in his example, so four observations are sufficient to
establish the parameters exactly (as long as the sets of X's don't happen to
be linearly dependent). The usual relationships hold if there are
uncertainties in the measures of the X's or in the parameters, so it is not
possible to get a highly reliable estimate of Y from highly unreliable
measures of the Xs. You get a little improvement in relative terms, and that
is all.

When I said that the experimenter would have to guess, I was thinking in
terms of the assumed form of the function. The implication that the observer
couldn't find the best values of the parameters was wrong, and your
objection is reasonable.

What on earth are you talking about? It seems to have no relevance to the
question of what the distribution of Y is after you have measured the X's.
My "scenario," as you call it, assumes that the experimenter measures the
X values in exactly the same way as when evaluating the correlations in the
first place. Who knows the "true" values? And what does it have to do with
the argument?

The question is whether we consider the sets of X's that Bruce presented to
be exact measures, or whether they represent attempts to measure variables
with a true value of 0 and a large amount of random variation. The former is
the case that Bruce considered; we have to treat Y as an exact function of
_each set of X's_ considered independently.

Of course, if Y is normally distributed with mean zero, then before you
gain any further information about it, your best bet is that any particular
measure will have a mean of zero. After you have gained some information
about it, by measuring one or more of the Xs, your best bet will be
some other value.

In Bruce's example, each X was a variable with a mean value of zero and a
certain distribution of values around zero. But he did not treat any X as a
random variable with a true value and a disturbance. That is how _you_ are
treating them. If that were the case, then it would not be true that _every_
set of four measures would yield a value of y exactly equal to the sum of
the four numbers. Instead, the relationship would be that the mean value of
Y is the sum of the mean values of the Xs, but there would be no reason to
suppose that the instantaneous value of Y would also be the sum of the
instantaneous values of the Xs -- nothing would constrain the individual
sets to sum to the same number. Bruce set it up so they DID sum to the
instantaneous values of Y -- that is why this was not a statistical problem,
but only one of solving simultaneous equations. The following shows that you
are considering this as a statistical problem (quite correctly, in terms of
that assumption):

Once you've done your regression analyses, and found that to your best
estimate Y = 1.5X1 + 2.3X2 + (-0.8)X3 + 0.15X4 + (random with sd 0.5), then
after you have measured X1 to be 2.0 on a specific occasion, your best
estimate of Y is 3.0, and it will have a standard error of estimate a
bit less than that of your original zero estimate.

This is perfectly right, but you would not expect EACH set of X's to fit the
equation exactly, as in Bruce's example.

Assuming that the
variance of each X is unity, I think that the variance around your estimate
of Y will drop from (1.5^2 + 2.3^2 + 0.8^2 + 0.15^2 + 0.5^2) to
(2.3^2 + 0.8^2 + 0.15^2 + 0.5^2). Could be wrong there, but I think not.

I think you've left out one of the X's -- X1, I think. If you include that,
the two sums are identical. The variance of the sum is the sum of the
variances, right? I don't know where you got the extra random variable, but
it's OK if you want it there. But this is not Bruce's example: the variance
of the X's is zero in his example. There is no measurement error.

-------------
I have no intention to be slippery, or to present anything misleading. My
objective, though it may sometimes seem otherwise, is clarity. In this
case, I used the special case of positive measures to avoid unnecessary
complication while making a general point.

But by doing so, you made a point that is valid only in this special case.
If you want to make a general point, you have to treat the general case.

If you want to do it right, you can work from the uncertainties of the
X values and the Y value before and after the measures are made...

There is no point in going on. We don't disagree about how to handle the
combining of uncertain measures. You have missed the point that in Bruce's
example, there are no uncertainties in any measures of X's, or in the values
of Y.

Best,

Bill P.

[From Bill Powers (970424.0915 MST)]

Hans Blom, 970424d--

You want to see behavior as what follows from (internal) (changes of)
reference levels. So you put the less certain and more abstract thing
(references) first and the more certain and more concrete things
(observable actions) second. Some people would say that this is a
logical error: explaining more certain things by less certain ones.
But I agree with you that this point of view has its uses.

I think you misinterpret me. Even in MCT, the same point I am making
applies. There is an observable (by an outside observer) input quantity that
you call x and I call qi. There is an observable disturbance, which we both
call d (although I am assuming here that Fd = 1*d. There is an observable
output that you call u and I call qo.

Suppose that an external observer failed to notice x (or qi), and
manipulated only d. In the case where the reference signal is constant, this
observer would see that a change in d produces an equal and opposite change
in u. However, this relationship is "equal and opposite" only in terms of
the effect on x; it may not be equal and opposite in terms of direct
measures of d or u (because, for example, there is a system function, or
what I call an environmental feedback function, intervening between u and
x). So it is unlikely that this observer will happen to make the
measurements in just the way needed to see the perfect symmetry.

Anyway, in this case the observer might well conclude that the control
system is sensing d (rather than x), and that the result of the sensing is
simply to cause a change in u, as if there were a straight-through
connection from d to u. Because of using inappropriate measures of d and u,
the observer will not see the almost-perfect negative correlation that
exists between them, or the nearly perfect cancellation of their effects on
x, but will find some lower correlation between them.

Also, the observer may find other d's that also appear to affect u via the
system. These other d's may enter in other places in the external system, or
they may be connected to x by other causal paths, yet they, too, will prove
to be correlated with the same measure of behavior, u. The correlation is
not likely to be high for the reasons already given, but they will quite
likely be nonzero -- in fact, the observer will discard them unless they are
nonzero.

Remember that this is all postulated on the assumption that the observer
does not realize that there is an x being controlled by the control system
through its output u, as is the case for both the MCT and the PCT models. In
the case of organisms, there is no nicely delineated "plant" that can be
seen in the environment -- there is just the whole environment, with the
organism acting in it and on it. Nothing stands out as the "plant" and no
variable is self-evidently a controlled x. There is no reason for the
observer to suspect, if the possibility of control has never been
considered, that d is not itself being sensed, and that u is not simply a
reaction to d -- a "stimulus."

So this is where the behavioral illusion comes into being, and how S-R
theory quite rationally, but mistakenly, arose. Overlooking the possibility
of control is what has made behavioral laws seem so uncertain and noisy. It
explains why physical variables that seem to have nothing to do with each
other nevertheless have similar or directly opposite effects on behavior --
on u. Why does opening a window cause someone to put on a sweater, while
building a fire in the fireplace causes him to take it off, and sunshine
coming in the window can make him take if off again? Without the realization
that this person is controlling his body temperature, these facts remain
mysterious. What does opening a window have in common with building a fire
or sunshine? And why should any of these "stimuli" cause the person to do
anything with a sweater?

But although I sympathize with your point of view, it has proven to
be extremely hazardous to redefine a term which has a certain well-
established connotation. Too often people seem to interpret the ":"
in B:CP as an "=", you yourself not excluded. It remains my personal
preference to use the word "behavior" for _observables_. In HPCT that
would be the collective outputs of the lowest hierarchical level.

I do use the term "behavior" for observables. More precisely, I like to use
the term "action", because behavior is often defined in terms of its effects
on the environment instead of in terms of actual outputs. The behavior we
observe organisms carrying out consists of the actions that are used in the
process of controlling perceptions. That is the expanded meaning of
"Behavior: the control of perception."

If we equate "behavior" with externally observable actions, we can
clearly see that in a control system "behavior" has _two_ "causes".
Take the theodolite controller: its output is determined by (1) its
internal reference level(s) _and_ (2) the external "disturbance".
When the disturbance suddenly goes to +10, the compensatory action/
output goes to -10. This is where stimulus-response theory applies --
_because a controller is at work_.

Yes, exactly. I agree. That is always how I use the term "behavior." You are
describing exactly the behavioral illusion. What is illusory about it is
that the stimulus seems to be directly sensed, and seems to cause the
response through some simple lineal chain of events. Most of the world of
psychology is ignorant of the true relationship (assuming, of course, that
PCT or MCT offers the correct explanation).

But changes of the reference level
translate into output changes as well. And stimulus-response theory
therefore posits "internal" stimuli (e.g. hunger or thirst) that the
observables can be attributed to. Not unreasonable: they are the
higher level references of an HPCT hierarchy.

Yes, that's true, too.

I think that as long as we keep in mind that externally visible
"behavior" has two causes, a fairly straightforward "translation"
between SR-theory and control theory is possible -- although control
theory gives additional insight. It is therefore, maybe, that PCT
impresses many people as "nothing but" something very much like a
top-down versus bottom-up coordinate transformation: useful, but
"nothing new". Alas: they miss the additional insight...

I'm very pleased with your comments here. All is not lost.

···

----------------------------------------
Incidentally, I used the wrong example of a disturbance in trying to make my
point about "reconstructing the disturbance" in your program. What I should
have done is this:

Inside the "disturbance" function in your program, declare a typed constant
(local variable) dd. Make dd := t, so dd increases linearly with time during
the run. Return the value,

disturbance := 10*sin(dd);

What your program does is estimate the _effect_ of dd on x, which is the
value returned by the disturbance function; what Martin is now calling the
disturbance signal. What I wanted to show is that the environmental variable
that actually produces this effect (dd) can be related to the effect through
any function at all. If the controlled variable is a light intensity, I can
disturb it by placing a lamp near the photocell. As I move the lamp closer
and farther way (so dd is measured as the radial distance of the lamp from
the photocell) the disturbing light intensity signal will vary as 1/dd^2.
The compensating change in u, in terms of its effect on x, will thus be
proportional to and opposite to 1/dd^2, not to dd. The linear correlation of
u vs dd will be considerably less than 1.00, even if the control system
operates perfectly and totally cancels the effect of the disturbance.

With dd being defined as in the program modification above, the output of
the control system (given a constant reference signal) will vary in a way
that cancels the _effect_ of dd on x. Because of the intervening
integrations representing the physical properties of the theodolite, the
changes in u will not simply be equal and opposite to these effects. And
since the effect of the disturbance is connected to dd through a sine
function, neither will the effects of the disturbance be directly related to
dd. So if we imagine manipulating dd and looking for correlations with u, it
is very hard to see how any correlation at all would be found. If we used a
damped mass on a spring as the load, there would be some proportional term
in the system equation, and then we would see some correlation, but still
far from 1.00 if the dynamic effects were comparable to the proportional
term. So we can see how failure to recognize the existence of control can
lead not only to a false picture of how behavior works, but to predictions
with an unnecessarily high amount of uncertainty.

Best,

Bill P.

[Martin Taylor 970425 10:35]

Bill Powers (970424.0915 MST)

In case it makes any difference, I think your message to Hans fully captures
my own understanding of the situation.

There is an observable (by an outside observer) input quantity that
you call x and I call qi. There is an observable disturbance, which we both
call d (although I am assuming here that Fd = 1*d. There is an observable
output that you call u and I call qo.

Suppose that an external observer failed to notice x (or qi), and
manipulated only d. In the case where the reference signal is constant, this
observer would see that a change in d produces an equal and opposite change
in u. However, this relationship is "equal and opposite" only in terms of
the effect on x; it may not be equal and opposite in terms of direct
measures of d or u (because, for example, there is a system function, or
what I call an environmental feedback function, intervening between u and
x). So it is unlikely that this observer will happen to make the
measurements in just the way needed to see the perfect symmetry.

And so on.

I think the core of this message should be submitted as one of the items in
the
introductory archives on the Web site, perhaps with trivial editing to remove
reference to MCT and to any other posting.

Martin

[From Bill Powers (970425.0853 MST)]

Martin Taylor 970425 10:35 --

]

Bill Powers (970424.0915 MST)

In case it makes any difference, I think your message to Hans fully
captures my own understanding of the situation.

Good, it's nice to agree and know what we're agreeing about.

Best,

Bill P.

[Martin Taylor 970424 13:30]

Bill Powers (970424.0751 MST)]

Martin Taylor 970423 10:00]--

>Why don't you accept it when I write it but are quite happy when Rick
>uses the same symbol for the same signal?

Because I know that the default meaning when Rick refers to the disturbance
is qd, whatever symbol he uses for it. Your default meaning is what you now
propose to call the disturbance signal. When Rick says "o + d" I
automatically translate, as he does, to "Fd(qd) + Fe(qo)", with the
understanding that the two functions are, for this simplified case only,
unity multipliers.

Well, I wish you would tell Rick that. He even wrote it explicitly as
p = f(o+d) the other day, and I had to correct him to the form you say
he implicitly uses. And to emphasize the point, he makes it very clear
that it is the summation of the effect on the input variable that is
important. He wouldn't say this, I suppose, if he didn't mean it.

And even this morning, he uses "disturbance" to mean not qd, but the
effect of the disturbing variable on qi, in a very elegant message.
I wish I could have written that, myself. But if I had, I suspect you
would have objected to the use of "disturbance" that way.

I ask only for the same default assumptions as you grant Rick--that he
means something sensible when he uses a term ambiguously, and not to
assume always that the non-sensible version must be what I mean.

···

-------------

But partitioning by some means is necessary if explicit use is to be made of
the information -- if the part of the perceptual signal magnitude due to the
disturbance is to be treated any differently from the part due to the
system's own action.

Maybe here's the nub of the problem. You may be thinking of "explicit
models" or some such. I'm not. I have no desire to partition the part of
the perceptual signal due to the disturbance from the part due to the
control unit's output. I'm talking a generalization of a correlational
analysis. There's no "explicit use" of the disturbance information as
distinct from the output information. What there is is a relationship,
some part of which I described but did not compute in a message a couple
of days ago about transport lag, some part of which is in the small residual
correlation between the disturbance and the perceptual signal.

The comparator compares the _whole_ perceptual signal
to the reference signal, not just the part due to the disturbance.

I do understand about superposition as a principle of linear analysis.

Yes, but in general, the linear analysis won't apply. If it would, there'd
be no advantage in even thinking about more general analyses, let alone
attempting them.

---------------------------
>On whether any information can be gained about a component from measurement
>of a sum, have you actually read Richard Kennaway's recent postings, which
>Rick praised so highly? Kennaway shows precisely how much information
>can be gained about a component from the sum. Using his formulae, I
>showed yesterday the rates for two signals having different correlations.
>
>I hoped you would take these numbers and consider them in the case where
>one signal is p and the other d (disturbance signal value). Instead, you
>choose to hold the self-contradictory position that Kennaway's paper and
>postings are fine, but his conclusions are not.

No, I simply accepted his computation of the number of bits, and understood
how a one-bit prediction amounts to the prediction of only the sign of a
relationship.

But for a continuous variable it doesn't even predict that, while at the
same time it actually predicts more. What it predicts is that measuring
one variable permits a halving of the standard deviation of the estimate
of the variable not measured. At the same time, for a continuous variable,
no number of bits can allow you to make an unambiguous prediction even of the
sign.

I haven't seen your computation of the number of bits
representing the disturbance signal as they relate to the information in the
perceptual signal.

You haven't? Did my Gilbert and Sullivan jibe upset you enough that you
didn't read to the end of that message? I presented a table that would be
part of it, the part you have in the past been interested in, and that
represents the degree to which control is imperfect. Some years ago Tom
Bourbon (I think) presented a set of data showing the correlations between
the disturbance (signal, it should have been, but I suppose it was "variable")
and the putative CCEV. If I remember, these hovered in the range of 0.2
or thereabouts. I've been trying to find them in the archives, with no
luck. If 0.2 is right for the central trend, my table seems to suggest
that the bit rate of a disturbance signal of 5 Hz bandwidth in the
perceptual signal (or equivalently, the mutual information being symmetric,
of the perceptual signal in the disturbance signal) is about 0.6 bits/sec.

This case is complicated seriously by the fact that the
flow of information through the input function is considerably reduced by
the negative feedback effects of the output -- which is only to repeat,
without much understanding, what you have said several times. If you come up
with a valid computation, I will probably accept that, too. See the above.

>If two measures have a correlation c, then each measure of one of them
>conveys -0.5*log2(1-c^2) bits about the other. If X = Y + Z, the correlation
>between X and Y depends on the variances of Y and Z in the way Richard
>described (if Y and Z are correlated, as they are in the case of o and d,
>then there's a cross product term that comes in as well, but that doesn't
>affect the relation between correlation and mutual information, at least
>for gaussian variables).

That's for the open-loop case, which is what Richard was discussing. The
closed-loop case, for all I know and for all you have demonstrated, may be
completely different. So get busy.

No, open and closed loop are merely mechanisms whereby correlations come to
be as they are. The correlations themselves map onto information transfer, or
rather, they provide lower bounds on the information transfer between the
correlated variables.

To see this latter point, imagine a relationship Y = |X|, with X having
a zero-centred normal distribution. Knowledge of X provides perfect
knowledge of Y (in info-theory terms, the mutual information of X and
Y equals the uncertainty of X) but the correlation between X and Y is zero.
(Knowledge of Y, on the other hand, leaves a 1-bit uncertainty about the
value of X, since the initial uncertainty associated with the probability
density distribution of Y is 1 bit larger than that of the pdd of X.)

>-------------
>In the same vein:
>>I think Richard Kennaway dealt with this interpretation very well.
>
>Yes. He did accept what I said, after all;-)
>
>>You're
>>saying that the experimenter measured the X's, and just happened to combine
>>them so that Y was predicted perfectly. If the law (known only to God) had
>>been Y = aX1 + bX2 + cX3 + dX4, the experimenter would have had to guess at
>>a, b, c, and c in order to get the perfect prediction. His guess that Y =
>>sum of X's would not predict very well at all.
>
>Bruce's example involved no guess about a, b, c, and d. A multiple
>regression analysis discovers an estimate of those values, with a precision
>that increases with the number of observations of the X's and Y. The rest
>of your comment on this matter is therefore invalid.

No, Bruce has now acknowledged that there is no uncertainty in the measures
of the X's or of Y in his example, so four observations are sufficient to
establish the parameters exactly (as long as the sets of X's don't happen to
be linearly dependent).

Well, he may have said he could measure the X's and the Y exactly (I don't
remember him doing so, but maybe he did). However, the example I provided
most explicitly did not make this assumption. The measurement error was
what I was getting at in the example, but you missed it, as your later
comments show.

I may have misinterpreted Bruce's description of his example. My intention
was to use his example as I thought he intended it, but making the measurement
error explicit.

When I said that the experimenter would have to guess, I was thinking in
terms of the assumed form of the function. The implication that the observer
couldn't find the best values of the parameters was wrong, and your
objection is reasonable.

And so is yours. The _forms_ possible are infinite in variety (literally--
there are aleph-2 of them). All we can do is hope that one of the more
mundane forms fits reasonably well, as best we can tell given the
measurement error we are stuck with. In the case in point, our minor God
was kind in designing a micro-universe in which our first guess was a good
one:-)

>Of course, if Y is normally distributed with mean zero, then before you
>gain any further information about it, your best bet is that any particular
>measure will have a mean of zero. After you have gained some information
>about it, by measuring one or more of the Xs, your best bet will be
>some other value.

In Bruce's example, each X was a variable with a mean value of zero and a
certain distribution of values around zero. But he did not treat any X as a
random variable with a true value and a disturbance. That is how _you_ are
treating them. If that were the case, then it would not be true that _every_
set of four measures would yield a value of y exactly equal to the sum of
the four numbers.

No, the measured value of Y would be the sum of the four Xs plus-or-minus
the vector sum of the measurement errors of the four Xs. We've been taking
those measurement errors to be equal, so the estimate of Y would have an
actual measurement standard deviation of half the sum of the standard
deviations of the Xs.

Instead, the relationship would be that the mean value of
Y is the sum of the mean values of the Xs, but there would be no reason to
suppose that the instantaneous value of Y would also be the sum of the
instantaneous values of the Xs

I don't understand this comment. In the world Bruce proposed, Y is ALWAYS
the sum of the four Xs. It's only our inability to measure the Xs and the
Y exactly that makes it not so in the measurements.

...Bruce set it up so they DID sum to the
instantaneous values of Y -- that is why this was not a statistical problem,
but only one of solving simultaneous equations.

I understood him differently, but there you are.

The following shows that you
are considering this as a statistical problem (quite correctly, in terms of
that assumption):

>Once you've done your regression analyses, and found that to your best
>estimate Y = 1.5X1 + 2.3X2 + (-0.8)X3 + 0.15X4 + (random with sd 0.5), then
>after you have measured X1 to be 2.0 on a specific occasion, your best
>estimate of Y is 3.0, and it will have a standard error of estimate a
>bit less than that of your original zero estimate.

This is perfectly right, but you would not expect EACH set of X's to fit the
equation exactly, as in Bruce's example.

Didn't you notice the term "+random with sd 0.5"? That's the degree to which
each set of measures would fail to predict the corresponding value of Y. The
reason it is there is that you can't measure the Xs exactly either when
finding the regression components or when measuring this instance of Y
(by way of the Xs).

>Assuming that the
>variance of each X is unity, I think that the variance around your estimate
>of Y will drop from (1.5^2 + 2.3^2 + 0.8^2 + 0.15^2 + 0.5^2) to
>(2.3^2 + 0.8^2 + 0.15^2 + 0.5^2). Could be wrong there, but I think not.

I think you've left out one of the X's -- X1, I think.

No I haven't. That's the one said was I measured in the sentences immediately
preceding your quote. The measurement of X1 is the reason why the variance
of Y is reduced.

If you include that,
the two sums are identical.

Sure, if I ignore the fact that I made a measurement of X1, things are as if I
hadn't made the measurement. What do you want me to conclude from that?

The variance of the sum is the sum of the
variances, right? I don't know where you got the extra random variable, but
it's OK if you want it there.

I got it from the measurement errors of the individual Xs.

You have missed the point that in Bruce's
example, there are no uncertainties in any measures of X's, or in the values
of Y.

Well, one of us missed Bruce's point. That's for sure.

Martin

[From Bill Powers (970425.1654 MST)]

Martin Taylor 970424 13:30 --

>Why don't you accept it when I write it but are quite happy when Rick
>uses the same symbol for the same signal?

Because I know that the default meaning when Rick refers to the
disturbance is qd, whatever symbol he uses for it.

Well, I wish you would tell Rick that. He even wrote it explicitly as
p = f(o+d) the other day,

That is quite correct. p, the perceptual signal, is the input function (f)
of the sum of Fd(qd) + Fe(qo), which with unity multipliers can be written
as (o + d) for short. Written out in full, his equation is

p = Fi[Fe(qo) + Fd(qd)]; compare with

p = f[ o + d ],

which holds only when Fe and Fd are unity multipliers. d always refers to qd.

If you want it in still more detail you can explicitly include the input
quantity qi:

qi = Fe(qo) + Fd(qd); [qi=qo+qd or o+d, with unity multipliers]

p = Fi(qi).

If you have the right default meaning of o and d in mind, there is no
difficulty with interpreting Rick's shorthand.

What Rick is talking about is that qi is sensed, not the component values
Fe(qo) and Fd(qd). You can conceptually divide qi into the part due to the
disturbance and the part due to the output, but the system itself perceives
only a function of the sum qi, represented by p. It is the physical signal p
that is compared with the reference signal, not just one conceptual
component of p.

The only way in which the control system itself could deduce that there is a
disturbance signal would be to contain a subsystem that SOMEHOW senses the
value of its own output quantity, the FORM OF THE FEEDBACK FUNCTION, and the
value of qi. It could then compute that

Fd(qd) = qi - Fe^-1(qo)

(which is basically how Hans' model deduces the disturbance).

And to emphasize the point, he makes it very clear
that it is the summation of the effect on the input variable that is
important. He wouldn't say this, I suppose, if he didn't mean it.

Yes, he did mean it, and for the reasons (I presume) that I have just outlined.

In order to make any use of information about the disturbance signal, the
system would first have to calculate the disturbance signal, and in order to
calculate that it would have to know the three items mentioned above: qi,
Fe, and qo. Without that knowledge, there would be nothing to be correlated
with the perception, and thus no way to know what part of the perception was
due to the disturbance signal. So without all that, there is no way the
system itself can make any use of knowledge about the disturbance signal.

You, as an external observer, can of course perform that calculation,
because you know qo, Fe, and qi -- you may even know, after the fact,
Fd(qd), or even Fd and qd separately. Thus you can calculate the correlation
between Fd(qd) and p, and derive a number showing how much of the variance
of p is due to Fd(qd), the disturbance signal. However, your knowledge of
this fact plays no part in the operation of the control system. The control
system controls WITHOUT knowing what part of p is due to the disturbance
signal and what part to qo. It does not know that because, in the PCT model,
there is no provision for obtaining the required knowledge of qo and Fe and
no machinery for computing qi - Fe^-1(qo).

Think about this, Martin. The reasoning you are using is exactly the
reasoning that leads to the MCT model. The MCT model does in fact perform
the calculations that allow computing the disturbance signal, and it
contains machinery (rather incompletely defined, to be sure) that allows the
control system itself to know the form of Fe. That is what the "world model"
is, the form of Fe. The MCT model _does_ make use of this information,
because it is designed to compute it explicitly. The system identification
methods implied in the background, and the adaptive Kalman filter methods,
are what generate the model of Fe, and they, too, are part of the MCT model.
All of that is needed, if your design is based on knowing the form of Fe and
calculating the disturbance signal. In THAT model, a calculation of the
amount of information in the perception about the disturbance signal is
perfectly relevant, because the system is creating and using that
information explicitly.

The great advantage of the PCT model is that it does NOT have to perform all
those calculations; it does not rely on knowing the value of Fd(qd)
separately, as the MCT model does. It does NOT have to know what part of p
is due to the disturbance signal and what part to Fe(qo). It needs no
knowledge of Fe, either. Adaptation can be carried out strictly on the basis
of the error signal; no model of the environment is needed (although as Hans
has said, the result is an _implicit_ model, or pseudo-model).

And even this morning, he uses "disturbance" to mean not qd, but the
effect of the disturbing variable on qi, in a very elegant message.

I, too, occasionally slip and say "disturbance" when I should have said
"effect of the disturbance". Your term, disturbance signal, will help in
avoiding that mistake, which I expect is motivated mostly by getting tired
of typing four words to designate a single variable.

Maybe here's the nub of the problem. You may be thinking of "explicit
models" or some such. I'm not. I have no desire to partition the part of
the perceptual signal due to the disturbance from the part due to the
control unit's output. I'm talking a generalization of a correlational
analysis.

There's no "explicit use" of the disturbance information as
distinct from the output information. What there is is a relationship,
some part of which I described but did not compute in a message a couple
of days ago about transport lag, some part of which is in the small
residual correlation between the disturbance and the perceptual signal.

In order to do a correlational analysis you need to know Fd(qd) explicitly.
......................

I haven't seen your computation of the number of bits
representing the disturbance signal as they relate to the information in the
perceptual signal.

You haven't? Did my Gilbert and Sullivan jibe upset you enough that you
didn't read to the end of that message? I presented a table that would be
part of it, the part you have in the past been interested in, and that
represents the degree to which control is imperfect.

But this table is good only for open-loop calculations. I suspect that
closed-loop calculations will provide some surprises.

Some years ago Tom
Bourbon (I think) presented a set of data showing the correlations between
the disturbance (signal, it should have been, but I suppose it was >"variable")
and the putative CCEV. If I remember, these hovered in the range of 0.2
or thereabouts.

In my various calculations like this, the correlation ranges between 0.2 and
-0.1; the average may be positive, but I haven't really studied that. In a
true integrating control system with noise in the system, I think the mean
correlation should approach 0.0.

That's for the open-loop case, which is what Richard was discussing. The
closed-loop case, for all I know and for all you have demonstrated, may
be completely different. So get busy.

No, open and closed loop are merely mechanisms whereby correlations come
to be as they are. The correlations themselves map onto information
transfer, or rather, they provide lower bounds on the information transfer
between the correlated variables.

That's what you say now, without having actually derived the correct
equations. You're extrapolating from experience with open-loop systems. In a
closed-loop system, variables that are treated as independent are not
actually independent, because of the feedback effects. And even I know that
certain assumptions about independence have to be met to make calculations
of correlation valid.

To see this latter point, imagine a relationship Y = |X|, with X having
a zero-centred normal distribution. Knowledge of X provides perfect
knowledge of Y (in info-theory terms, the mutual information of X and
Y equals the uncertainty of X) but the correlation between X and Y is >zero.

But what if Y = |X| while X = k*(|Y0 - |Y}||)?

There's a whole lot more in your post, but I'm not up to dealing with it now.

I think the most important point where we might get somewhere is in the
above discussion of MCT. Without knowledge of the MCT model, I couldn't have
said it that way. Maybe this will get us somewhere -- anyway, I'm pooped.

Best,

Bill P.

[Martin Taylor 970426 21:45]

Bill Powers (970425.1654 MST)

I think that maybe, just maybe, I'm getting a glimmering of a possible
insight into why we have this long-standing block in our communication
in respect of information. Let's give it a go.

What Rick is talking about is that qi is sensed, not the component values
Fe(qo) and Fd(qd). You can conceptually divide qi into the part due to the
disturbance and the part due to the output, but the system itself perceives
only a function of the sum qi, represented by p. It is the physical signal p
that is compared with the reference signal, not just one conceptual
component of p.

The only way in which the control system itself could deduce that there is a
disturbance signal would be to contain a subsystem that SOMEHOW senses the
value of its own output quantity, the FORM OF THE FEEDBACK FUNCTION, and the
value of qi.

OK. From this I deduce that you think that the informational analysis in
some way presupposes that the control system separates out _two_ input
signals, one based on the disturbance signal and one on the output signal.
Is that correct?

If you are supposing this, then unsuppose it. The analysis needs no more
wires, variables, or signal lines than are in the normal everyday diagram,
and "p" is still a scalar variable that is a function of qi.

I won't go further along that line until I know whether my surmise is true.
Nothing I have said should have been construed to state that the control
system _does_ separate the component of the input due to the disturbance
from that due to the output. I did demonstrate once that it is _possible_
to reconstruct the disturbance given no variables except the qi waveform,
using the fixed function that happened to be Fe(Fo()). That demonstration
only showed that the perceptual signal does carry a lot of information
about the disturbance, not that this information is segregated in the
action of the control system.

Note that this is NOT what Hans used. He used Fe(Fo(o)), which I consider
cheating. And it's not what Rick (Rick Marken(970425.1730) commented on:

Of course, I am going with Hans' assumption that g(d) = 1*d. I know, in
other words, that if p = f(u) + g(d) then you would have to have f(u)
_and_ g(d) (not just d) to solve for p.

In those terms, what is needed is f() and g(d), not f(u) and g(d). There's
a world of difference. But it's rather irrelevant that an analyst _can_
do this kind of partitioning, except to show that the information about
the disturbance signal in the perceptual signal is not lost--at least the
part that has actually contributed to control is not. The rest shows up
as the correlation between the disturbance signal and the perceptual signal.

···

------------------
On correlation and bits

I presented a table that would be
part of it, the part you have in the past been interested in, and that
represents the degree to which control is imperfect.

But this table is good only for open-loop calculations. I suspect that
closed-loop calculations will provide some surprises.

If you look at it, you'll see that it has nothing to do with open or
closed loop calculations. It has only to do with the values of data.
The value of the bit rate calculated from the correlation is correct
for bivariate Gaussian distributions, but for other distributions
that calculation understimates the bit rate (I'm pretty sure but not
certain).

No, open and closed loop are merely mechanisms whereby correlations come
to be as they are. The correlations themselves map onto information
transfer, or rather, they provide lower bounds on the information transfer
between the correlated variables.

That's what you say now, without having actually derived the correct
equations.

Richard Kennaway did that. The correct equations have to do with what
value of x occured along with what value of y on this occasion, that
occasion, and the other occasion(s). There's no issue of how the values
were produced, just what the values turned out to be. You can calculate
the correlation between the number of hairs on a dog and the number of
sunspots on the visible part of the sun when it was born, if you want.

You're extrapolating from experience with open-loop systems. In a
closed-loop system, variables that are treated as independent are not
actually independent, because of the feedback effects. And even I know that
certain assumptions about independence have to be met to make calculations
of correlation valid.

You don't _assume_ independence, you look to see if it is there.

Calculations of correlation are done to determine whether there is
_linear_ independence between the variables. You are presented with a set
of data pairs ({x1,y1}, {x2,y2},....{xn,yn}) and you calculate a correlation.
Or you can calculate a mutual information measure. Or you can assume that
the pairs come from a bivariate Gaussian distribution and compute a
lower-bound bit rate from the correlation. What assumptions were you
thinking of?

To see this latter point, imagine a relationship Y = |X|, with X having
a zero-centred normal distribution. Knowledge of X provides perfect
knowledge of Y (in info-theory terms, the mutual information of X and
Y equals the uncertainty of X) but the correlation between X and Y is
zero.

But what if Y = |X| while X = k*(|Y0 - |Y}||)?

There's some problem with your notation here, but it doesn't matter,
provided that whatever you intended is logically consistent with the
specifications, namely Y = |X| and X is from a zero-centred distribution.

If X is symmetrically distributed about zero, then the correlation between
X and Y will be zero, while the mutual information between X and Y is
the uncertainty of Y (I said "the uncertainty of X" before, but that's
an obvious mistake, like my far-too-common mixup of plus and minus signs
when I do algebra:-(. X clearly has one bit more uncertainty than Y, and
it is Y that is completely determined by X). It doesn't matter at all
if there are other relationships that are logically consistent with the
specified one. Y is still fully specified by X, but X has a one-bit
uncertainty after Y has been specified (unless Y=0 exactly:-).

What function _did_ you mean, by the way? I'm guessing that it was
X = k*(|Y0|-|Y|). Let's try and see if this is consistent with the
specification of the situation, which is: Y = |X| and X is symmetrically
distributed about zero.

Substituting for Y in your expression, we have X = k*(|Y0|-|X|).
Collecting, we get X - k*|X| = k*|Y0|. Assuming Y0 and k to be fixed,
this has only one solution for X, doesn't it? And that's inconsistent
with the specification that X is symmetrically distributed about zero
(in the original, it was more precise, a zero-centred Gaussian), so this
formula is invalid.

When we are dealing with data obtained in a closed loop situation, that's
where the data came from. Interesting, perhaps, to someone who might like
to guess what the correlation implies, but not to someone who just wants
to compute the correlation (or the mutual information) between the variables.
Such a person needs only the data--just as the control system needs only
qi, and knows nothing of the cause of any disturbance.

Martin

[From Bill Powers (970426.2210 MST)]

Martin Taylor 970426 21:45--

I think that maybe, just maybe, I'm getting a glimmering of a possible
insight into why we have this long-standing block in our communication
in respect of information. Let's give it a go.

OK. From this I deduce that you think that the informational analysis in
some way presupposes that the control system separates out _two_ input
signals, one based on the disturbance signal and one on the output signal.
Is that correct?

Yes.

If you are supposing this, then unsuppose it. The analysis needs no more
wires, variables, or signal lines than are in the normal everyday diagram,
and "p" is still a scalar variable that is a function of qi.

But you assume that there is available some measure of the disturbance
signal other than qi itself. For the analyst this is true -- the analyst
knows qd and Fd, or can at least assume an equivalent qd and Fd. But for the
control system it is not true: the control system knows only qi.

I won't go further along that line until I know whether my surmise is >true.
Nothing I have said should have been construed to state that the control
system _does_ separate the component of the input due to the disturbance
from that due to the output. I did demonstrate once that it is _possible_
to reconstruct the disturbance given no variables except the qi waveform,
using the fixed function that happened to be Fe(Fo()). That demonstration
only showed that the perceptual signal does carry a lot of information
about the disturbance, not that this information is segregated in the
action of the control system.

But you have also said that the better the control, the less information the
perceptual signal carries about the disturbance signal, with a limiting
amount of 0 for perfect control. This is why I keep saying that you have not
yet really solved this problem: what you are saying in words, and mostly
from personal intuition, needs to be formally derived in mathematics. If the
amount of information in p about the disturbance signal is a function of the
quality of control, you need to prove that and quantify the statement.

And by the way, since we have agreed that the default meaning of disturbance
is qd, I do wish you would use your own term and consistently say
disturbance _signal_ instead of just disturbance. The two terms mean
different things.

Given all the functions and variables in the loop (qi, Fi, comparator, Fo,
and Fe), plus the reference signal, it is possible to compute qo. Once you
know qo, you can deduce the part of qi that is due to qo, the remainder
presumably being due to some disturbance signal. But there is no independent
way to determine the disturbance signal, unless you know qd and Fd.

In {Rick's] terms, what is needed is f() and g(d), not f(u) and g(d).

Yes, because u (qo) can be calculated from qi, Fi, comparator, r, Fo, and Fe.

But it's rather irrelevant that an analyst _can_
do this kind of partitioning, except to show that the information about
the disturbance signal in the perceptual signal is not lost--at least the
part that has actually contributed to control is not. The rest shows up
as the correlation between the disturbance signal and the perceptual >signal.

You are speaking from intuition here; I repeat, your statements won't mean
anything until you have verified them by a mathematical derivation.

------------------
On correlation and bits

I presented a table that would be
part of it, the part you have in the past been interested in, and that
represents the degree to which control is imperfect.

But this table is good only for open-loop calculations. I suspect that
closed-loop calculations will provide some surprises.

If you look at it, you'll see that it has nothing to do with open or
closed loop calculations. It has only to do with the values of data.
The value of the bit rate calculated from the correlation is correct
for bivariate Gaussian distributions, but for other distributions
that calculation understimates the bit rate (I'm pretty sure but not
certain).

But if y = f(x,y), the correlation you calculate between y and f(x,y) is not
the correlation between x and y. You must first separate the variables.
You can calculate such a correlation but it is not the true correlation
between x and y. If the data are obtained in a system where y = f(x,y), the
correlation you measure is not the correlation between x and y.

No, open and closed loop are merely mechanisms whereby correlations come
to be as they are. The correlations themselves map onto information
transfer, or rather, they provide lower bounds on the information
transfer between the correlated variables.

That's what you say now, without having actually derived the correct
equations.

Richard Kennaway did that. The correct equations have to do with what
value of x occured along with what value of y on this occasion, that
occasion, and the other occasion(s). There's no issue of how the values
were produced, just what the values turned out to be. You can calculate
the correlation between the number of hairs on a dog and the number of
sunspots on the visible part of the sun when it was born, if you want.

There's a difference between calculating the correlation in a data set
consisting of pairs of numbers, and calculating it based on a theoretical
model. If you see that y = 10*x, you would expect a correlation of 1.0
between x and y. But if, unknown to you, x = y^3, the actual theoretical
correlation, and the correlation you will measure, will be zero, because the
solution is y = 100, a constant.

You're extrapolating from experience with open-loop systems. In a
closed-loop system, variables that are treated as independent are not
actually independent, because of the feedback effects. And even I know
that certain assumptions about independence have to be met to make
calculations of correlation valid.

You don't _assume_ independence, you look to see if it is there.

You can do that only if you have experimental data to look at. When you're
trying to do theoretical calculations based on relationships in a model, you
must start by showing that the variables you want to correlate are not,
through some pathway other than the one you're considering, related to each
other.

Calculations of correlation are done to determine whether there is
_linear_ independence between the variables. You are presented with a set
of data pairs ({x1,y1}, {x2,y2},....{xn,yn}) and you calculate a >correlation.
Or you can calculate a mutual information measure. Or you can assume that
the pairs come from a bivariate Gaussian distribution and compute a
lower-bound bit rate from the correlation. What assumptions were you
thinking of?

I was thinking of theoretical relationships, such as in your example of
Y=|X|, as in the following paragraph. If X depends on Y by some other
pathway, so there is a closed loop, the theoretical correlation is no longer
the same.

To see this latter point, imagine a relationship Y = |X|, with X having
a zero-centred normal distribution. Knowledge of X provides perfect
knowledge of Y (in info-theory terms, the mutual information of X and
Y equals the uncertainty of X) but the correlation between X and Y is
zero.

But what if Y = |X| while X = k*(|Y0 - |Y}||)?

There's some problem with your notation here, but it doesn't matter,
provided that whatever you intended is logically consistent with the
specifications, namely Y = |X| and X is from a zero-centred distribution.

I was trying to introduce a second relationship that creates a closed loop.
If X depends on Y at the same time that Y = |X|, your statements about
"perfect knowledge" and so forth do not hold.

Let's try and see if this is consistent with the
specification of the situation, which is: Y = |X| and X is symmetrically
distributed about zero.

Substituting for Y in your expression, we have X = k*(|Y0|-|X|).
Collecting, we get X - k*|X| = k*|Y0|. Assuming Y0 and k to be fixed,
this has only one solution for X, doesn't it? And that's inconsistent
with the specification that X is symmetrically distributed about zero
(in the original, it was more precise, a zero-centred Gaussian), so this
formula is invalid.

The formula is valid if, in fact, there is a feedback path from Y to X; in
that case it must be your assumption that X is symmetrically distributed
around zero that is inconsistent with the possible behavior of the system,
or else we are seeing a degenerate case (the symmetrical distribution has a
standard deviation of zero).

When we are dealing with data obtained in a closed loop situation, that's
where the data came from. Interesting, perhaps, to someone who might like
to guess what the correlation implies, but not to someone who just wants
to compute the correlation (or the mutual information) between the
variables. Such a person needs only the data--just as the control system
needs only qi, and knows nothing of the cause of any disturbance.

That is certainly true if we are dealing with DATA. But you are starting
with a theoretical model, and assuming what the data will look like based on
a theoretical analysis. It is quite possible that if you predict
theoretically what various correlations will be, but without doing anything
special to take feedback into account, the data would not be what you expect.

I could certainly be wrong about that, but so could you. We will not know
the truth until you have produced a rigorous mathematical analysis of
statistical relationships in a closed-loop system. I don't believe you can
do that based on intuitions derived from experience with open-loop
relations. You say that there is no difference, but that is exactly what
remains to be demonstrated.

Best,

Bill P.

[Martin Taylor 970427 17:05]

Bill Powers (970426.2210 MST)]

We reallyt must live on different planets. I can't make head nor tail
of half of your posting. But we forge ahead, regardless, through the murk

(Incidentally, I won't be able to do much more of this before leaving
for a month at the end of next week. I've suddenly been told I have
to prepare a presentation, and I haven't even thought about the material
yet.)

Martin Taylor 970426 21:45--

I deduce that you think that the informational analysis in
some way presupposes that the control system separates out _two_ input
signals, one based on the disturbance signal and one on the output signal.
Is that correct?

Yes.

If you are supposing this, then unsuppose it. The analysis needs no more
wires, variables, or signal lines than are in the normal everyday diagram,
and "p" is still a scalar variable that is a function of qi.

But you assume that there is available some measure of the disturbance
signal other than qi itself. For the analyst this is true -- the analyst
knows qd and Fd, or can at least assume an equivalent qd and Fd. But for the
control system it is not true: the control system knows only qi.

That's correct. When you are analyzing a control system, you are an analyst,
aren't you? You want to know, for example, how much the action of the
control system has reduced the effect the disturbing variable would otherwise
have had on the input variable. You may want to know other things, such
as the spectrum of the error signal. The control system doesn't "know"
about that, does it? Nor does it "know" qi. The analyst knows qi, and
the analyst knows p, the perceptual signal that is a function of qi. The
analyst knows the function, the control system doesn't. The control system
only behaves the way it behaves, having the functions it has, and being
provided with the inputs it is provided with--the reference signal and
the disturbance signal.

I did demonstrate once that it is _possible_
to reconstruct the disturbance given no variables except the qi waveform,
using the fixed function that happened to be Fe(Fo()). That demonstration
only showed that the perceptual signal does carry a lot of information
about the disturbance, not that this information is segregated in the
action of the control system.

But you have also said that the better the control, the less information the
perceptual signal carries about the disturbance signal, with a limiting
amount of 0 for perfect control. This is why I keep saying that you have not
yet really solved this problem: what you are saying in words, and mostly
from personal intuition, needs to be formally derived in mathematics. If the
amount of information in p about the disturbance signal is a function of the
quality of control, you need to prove that and quantify the statement.

What is your measure of the quality of control? There are quite a few, I
believe, but they all come down to the reduction in influence the disturbing
variable has on the perceptual signal. Informationally, that concept is
equivalent to reduction of the mutual information between the perceptual
signal and the disturbance signal.

And by the way, since we have agreed that the default meaning of disturbance
is qd, I do wish you would use your own term and consistently say
disturbance _signal_ instead of just disturbance. The two terms mean
different things.

I try, but when the meaning is obvious, but it's hard to keep using the long
form--rather like forgetting to say "daily newspaper" and using "paper"
instead, when "paper" is already used for something else. But I'll
continue to try.

Given all the functions and variables in the loop (qi, Fi, comparator, Fo,
and Fe), plus the reference signal, it is possible to compute qo. Once you
know qo, you can deduce the part of qi that is due to qo, the remainder
presumably being due to some disturbance signal. But there is no independent
way to determine the disturbance signal, unless you know qd and Fd.

In {Rick's] terms, what is needed is f() and g(d), not f(u) and g(d).

Yes, because u (qo) can be calculated from qi, Fi, comparator, r, Fo, and Fe.

That's right. That's why it works. The point of the demo was that there
was no VARIABLE RELATED TO g(d) other than the perceptual signal, and
nevertheless the fluctuations in g(d) could be reconsituted pretty accurately.
The quality of the reconstitution is the same as the quality of control.
This proves that the fluctuations in the perceptual signal convey all
the information from the disturbance signal that is used in control. The
fact that to make the reconstruction requires various functions and a variable
that is independent of g(d) in no way alters that proof. They don't have
any fluctuations related to g(d). Only the perceptual signal does.

But it's rather irrelevant that an analyst _can_
do this kind of partitioning, except to show that the information about
the disturbance signal in the perceptual signal is not lost--at least the
part that has actually contributed to control is not. The rest shows up
as the correlation between the disturbance signal and the perceptual >signal.

You are speaking from intuition here; I repeat, your statements won't mean
anything until you have verified them by a mathematical derivation.

The last sentence might be constured as "intuition". But even _you_ can't
assert that a high correlation between the perceptual and disturbance signals
doesn't show failure of control. Or can you?...I sometimes wonder just what
strange things you _can_ say about correlation.

------------------
On correlation and bits

I presented a table that would be
part of it, the part you have in the past been interested in, and that
represents the degree to which control is imperfect.

But this table is good only for open-loop calculations. I suspect that
closed-loop calculations will provide some surprises.

If you look at it, you'll see that it has nothing to do with open or
closed loop calculations. It has only to do with the values of data.
The value of the bit rate calculated from the correlation is correct
for bivariate Gaussian distributions, but for other distributions
that calculation understimates the bit rate (I'm pretty sure but not
certain).

But if y = f(x,y), the correlation you calculate between y and f(x,y) is not
the correlation between x and y.

Agreed. The correlation between y and f(x,y) is 1.0.

... open and closed loop are merely mechanisms whereby correlations come
to be as they are. The correlations themselves map onto information
transfer, or rather, they provide lower bounds on the information
transfer between the correlated variables.

That's what you say now, without having actually derived the correct
equations.

Richard Kennaway did that. The correct equations have to do with what
value of x occured along with what value of y on this occasion, that
occasion, and the other occasion(s). There's no issue of how the values
were produced, just what the values turned out to be. You can calculate
the correlation between the number of hairs on a dog and the number of
sunspots on the visible part of the sun when it was born, if you want.

There's a difference between calculating the correlation in a data set
consisting of pairs of numbers, and calculating it based on a theoretical
model.

There's certainly a difference between predicting a certain value of
correlation and finding it in the data. That I'll grant. But what does
this have to do with anything?

If you see that y = 10*x, you would expect a correlation of 1.0
between x and y. But if, unknown to you, x = y^3, the actual theoretical
correlation, and the correlation you will measure, will be zero, because the
solution is y = 100, a constant.

??? This makes absolutely no sense to me. It's a word jumble.

Calculations of correlation are done to determine whether there is
_linear_ independence between the variables. You are presented with a set
of data pairs ({x1,y1}, {x2,y2},....{xn,yn}) and you calculate a

correlation.

Or you can calculate a mutual information measure. Or you can assume that
the pairs come from a bivariate Gaussian distribution and compute a
lower-bound bit rate from the correlation. What assumptions were you
thinking of?

I was thinking of theoretical relationships, such as in your example of
Y=|X|, as in the following paragraph. If X depends on Y by some other
pathway, so there is a closed loop, the theoretical correlation is no longer
the same.

If Y = |X|, then Y = |X|. It doesn't matter what else is true, it is true
that no matter what value you measure for X, the value of Y is its absolute
magnitude. You can't even _get_ a correlation until you have measured
several different values of both. We know from the specification that
X is symmetrically distributed around zero. That specification ensures
that there is zero correlation between X and Y, even though knowing X
enables you to know Y.

To see this latter point, imagine a relationship Y = |X|, with X having
a zero-centred normal distribution. Knowledge of X provides perfect
knowledge of Y (in info-theory terms, the mutual information of X and
Y equals the uncertainty of X) but the correlation between X and Y is
zero.

But what if Y = |X| while X = k*(|Y0 - |Y}||)?

There's some problem with your notation here, but it doesn't matter,
provided that whatever you intended is logically consistent with the
specifications, namely Y = |X| and X is from a zero-centred distribution.

I was trying to introduce a second relationship that creates a closed loop.
If X depends on Y at the same time that Y = |X|, your statements about
"perfect knowledge" and so forth do not hold.

Of course they hold. Each time you get a value of X, you find that the value
of Y is X if the measured X turns out to be positive, and is -X if the
measured X turns out to be negative. That's perfect information about Y
given X.

Let's try and see if this is consistent with the
specification of the situation, which is: Y = |X| and X is symmetrically
distributed about zero.

Substituting for Y in your expression, we have X = k*(|Y0|-|X|).
Collecting, we get X - k*|X| = k*|Y0|. Assuming Y0 and k to be fixed,
this has only one solution for X, doesn't it? And that's inconsistent
with the specification that X is symmetrically distributed about zero
(in the original, it was more precise, a zero-centred Gaussian), so this
formula is invalid.

The formula is valid if, in fact, there is a feedback path from Y to X;

Well, there can't be, can there, if to have one is inconsistent with the
specifications. But I wouldn't be surprised to find you _could_ construct
a feedback path that would allow the demonstration data to exist. I leave
that as an exercise for the reader. It doesn't interest me.

in
that case it must be your assumption that X is symmetrically distributed
around zero that is inconsistent with the possible behavior of the system,
or else we are seeing a degenerate case (the symmetrical distribution has a
standard deviation of zero).

You are really strange. I provide an example of data in which a zero
correlation goes along with high mutual information, to show that the
correlation measure is only a lower bound on I(X|Y). You construct
a circuit in which the demonstration data won't occur, and use it to
argue that therefore the arithmetic is wrong. I don't understand this
mode of argument, unless it is a rhetorical trick to ensure that the
point is lost in the fog. If that's what it is, I understand it and don't
like it.

If you prefer, I here is a list of data in which Y = |X| and ask you to
compute the correlation. You can construct whatever circuit you like that
might have generated those data, but if your circuit wouldn't have generated
them, it will be an irrelevant circuit.

X = 1, 3, -3, -2, 1, 4, -1, -1, 2, -4
Y = 1, 3, 3, 2, 1, 4, 1, 1, 2, 4

Read Section 5 of Kennaway's paper, if it would help.

Here's another point to mull over. If X provides B bits of information about
Y, Y provides B bits about X. It matters not a whit whether X influences Y,
X and Y are both influenced by Z, Y influences X, or none of these are true.
Information, like correlation, carries no implication of causality.

Martin

[From Bill Powers (970427.2009 MST)]

Agreed. The correlation between y and f(x,y) is 1.0.

Yes, but the correlation between x and y is not necessarily 1.0.
...

If Y = |X|, then Y = |X|.

Ah, I see the problem.

Suppose you observe that in an operant conditioning cage, the rate of
reinforcement R is the rate of behavior B divided by the ratio m: R = B/m.

From this you can calculate that B = mR. However, this does not mean that if

you cause reinforcements to appear at the rate R, the behaviors will appear
at the rate mR. This is a _unidirectional_ relationship. While the algebra
says that you can express either variable in terms of the other, it doesn't
follow that the reverse calculation has any physical meaning. B may be
fraught with information about R, but R contains no information about B,
because arbitrarily changing R will have no effect on B.

Maybe this unidirectionality has already been taken into account in the
informational calculations, but I don't recall the subject having been
brought up before. Anyway, this is what I was getting at.

You are really strange. I provide an example of data in which a zero
correlation goes along with high mutual information, to show that the
correlation measure is only a lower bound on I(X|Y). You construct
a circuit in which the demonstration data won't occur, and use it to
argue that therefore the arithmetic is wrong. I don't understand this
mode of argument, unless it is a rhetorical trick to ensure that the
point is lost in the fog. If that's what it is, I understand it and don't
like it.

My point was that if you see that Y depends on X through some physical path,
you may predict a correlation between X and Y, but if this (as I can now
say) is a _unidirectional_ relationship, it is possible that X depends on Y
in another fashion. If that is so, then you could find, as in my example,
that the two variables must actually be constant, so no correlation could be
calculated or observed.

I see now that this wouldn't necessarily change the correlation of X with Y
(where the result is not a constant value), but it would certainly affect
the flow of information. You could _calculate_ the "mutual" information only
if the relationship were bidirectional. Otherwise your calculation would
indicate only the flow in one direction.

If you prefer, I here is a list of data in which Y = |X| and ask you to
compute the correlation. You can construct whatever circuit you like that
might have generated those data, but if your circuit wouldn't have generated
them, it will be an irrelevant circuit.

X = 1, 3, -3, -2, 1, 4, -1, -1, 2, -4
Y = 1, 3, 3, 2, 1, 4, 1, 1, 2, 4

Read Section 5 of Kennaway's paper, if it would help.

My point was that you can generate data of this sort and calculate the
correlation, but when you claim that it might be produced by a system you're
analyzing theoretically, you could be wrong: it might be impossible for this
set of data to be produced by the actual system, because you have left out
the reverse relationship. I gave an example of this. I see now that this
example is much more of a special case than I thought when I gave it, and of
course without introducing unidirectionality, my example would make no
sense. I can see why you were puzzled, when I claimed that if x = y, y does
not necessarily equal x. I hope this makes more sense when you understand
that I was speaking of unidirectional relationships. If you have y = x in
the direction from x to y, but x = y/3 in the other direction, the net
result is that y = x = 1/3.

Maybe what all this boils down to is rather simple. If you have two
_unidirection_ relationships in a feedback loop, constraints are put on the
possible values of x and y that are not evident when you consider only one
of the relationships, and treat it as bidirectional.

Here's another point to mull over. If X provides B bits of information
about Y, Y provides B bits about X. It matters not a whit whether X
influences Y, X and Y are both influenced by Z, Y influences X, or none of
these are true. Information, like correlation, carries no implication of
causality.

Yes, I can accept that. Y provides B bits of information about X, but this
does not mean that by manipulating Y you can induce or predict changes in X
-- unless the function is bidirectional.

From the standpoint of the analyst, this may be irrelevant. I wouldn't know.

Best,

Bill P.

[Martin Taylor 970427 13:30]

Bill Powers (970427.2009 MST)]

Agreed. The correlation between y and f(x,y) is 1.0.

Yes, but the correlation between x and y is not necessarily 1.0.

No. It could be anything.

...

If Y = |X|, then Y = |X|.

Ah, I see the problem.

Suppose you observe that in an operant conditioning cage, the rate of
reinforcement R is the rate of behavior B divided by the ratio m: R = B/m.

From this you can calculate that B = mR. However, this does not mean that if

you cause reinforcements to appear at the rate R, the behaviors will appear
at the rate mR.

No, nothing in correlational or information analysis says it will. What
either of them say is something like: under the conditions in which the
data were obtained, the correlation was 0.73, or the mutual information
was 0.6 bits/datum.

This is a _unidirectional_ relationship. While the algebra
says that you can express either variable in terms of the other, it doesn't
follow that the reverse calculation has any physical meaning.

No. The _direct_ relationship also may have no physical meaning (and
often doesn't).

B may be
fraught with information about R, but R contains no information about B,
because arbitrarily changing R will have no effect on B.

This is wrong, because if B has information about R, then R has the
same amount of information about B.

Maybe this unidirectionality has already been taken into account in the
informational calculations, but I don't recall the subject having been
brought up before. Anyway, this is what I was getting at.

Yes, I see. You were thinking of the correlational or informational
analysis as having a direction and a mechanism, whereas either is just
a measure like "variance", which can be computed from any paired datasets.
It's probable that if two variables are correlated, somewhere in the
background there is a variable that influences both, or perhaps one
influences the other. But it's not even guaranteed that there is a
common influencing variable (hard to see how the correlation could be
sustained otherwise, though).

My point was that if you see that Y depends on X through some physical path,
you may predict a correlation between X and Y, but if this (as I can now
say) is a _unidirectional_ relationship, it is possible that X depends on Y
in another fashion. If that is so, then you could find, as in my example,
that the two variables must actually be constant, so no correlation could be
calculated or observed.

Fine, but you are going beyond the example, now. If you predict a correlation
from your model, and the correlation doesn't happen, then your model is
probably wrong. If you predict that measurements will vary, and they don't,
your model is probably wrong. But in the example, there was no such model--
just a description of some hypothetical data that would show high
mutual information and zero correlation.

I see now that this wouldn't necessarily change the correlation of X with Y
(where the result is not a constant value), but it would certainly affect
the flow of information. You could _calculate_ the "mutual" information only
if the relationship were bidirectional. Otherwise your calculation would
indicate only the flow in one direction.

No, there's NO "flow" of information implied. Any "flow" comes from the
model that you produce to explain the observations. The observations aren't
biased as to whether X influences Y, Y influences X, or Z influences both
independently. Each way, for the same measured set of values of X and Y,
the mutual information is the same, and you learn as much about X from
measuring Y as you do about Y from measuring X.

Algebraically, if the uncertainty of X is written U(X), and the uncertainty
of the joint distribution of X and Y is written U(X,Y), then the mutual
information U(X:Y) = U(X) + U(Y) - U(X,Y). It's symmetric in X and Y. Also,
U(X) = U(X:Y) + Uy(X)
U(Y) = U(X:Y) + Ux(Y)

where I used a small x or y for a subscript in Ux(Y); it means the uncertainty
of X when a particular value of y is known, weighted by the probability of
getting that value of y.

My point was that you can generate data of this sort and calculate the
correlation, but when you claim that it might be produced by a system you're
analyzing theoretically, you could be wrong: it might be impossible for this
set of data to be produced by the actual system, because you have left out
the reverse relationship.

Yes, the system one proposes may not actually produce the data one supposed
it would. Then one's analysis of the system would have been wrong. But this
is starting from a system, and hypothesizing that it would produce a
certain kind of data. That's what we do when we want to compare theory
with practice. It's the reverse of what I was doing, which wqas hypothesizing
data that has certain properties, with no suggestion of mechanism.

Maybe what all this boils down to is rather simple. If you have two
_unidirection_ relationships in a feedback loop, constraints are put on the
possible values of x and y that are not evident when you consider only one
of the relationships, and treat it as bidirectional.

No question about it. That's quite true.

Here's another point to mull over. If X provides B bits of information
about Y, Y provides B bits about X. It matters not a whit whether X
influences Y, X and Y are both influenced by Z, Y influences X, or none of
these are true. Information, like correlation, carries no implication of
causality.

Yes, I can accept that. Y provides B bits of information about X, but this
does not mean that by manipulating Y you can induce or predict changes in X
-- unless the function is bidirectional.

No, nor can you cause changes in X by manipulating Y if there's no
influential relationship between them. But the informational relationship
can still hold perfectly well. In the situation analyzed, you may or may
not have been manipulating one or the other. You might have been just
observing. If that was the case, you have changed the system by performing
the manipulation--you've changed the influences, and therefore possibly
the correlation.

Consider this example to see what I mean. You measure the brightness of
moonlight at midnight every night at some place, and also the level of
the water in a tidal harbour six hours later somewhere else. You notice a
correlation, such that when the light is bright the water is high. (This is
Arizona, so there are no clouds to confuse the measurement:-) You need more
light, so you say "Aha. I'll just block the harbour entrance and pump in
more water." Then you find that doesn't work, so you say "Hmm. I can't
get more light, but perhaps I can stop the water from coming up so high"
so you shade the light meter. But the water still rises in the harbour.

In this example, you _don't_ find the correlation if you do the manipulation.

The light and the water height at midnight are both related to the phase
of the moon (and other things, so the correlation isn't perfect). You
don't change that fact by your manipulation, but you do change the efficacy
of the moon's influence by either manipulation. So long as you don't
manipulate, you can judge the midnight water height roughly by measuring the
moonlight, and you can judge the moonlight roughly by measuring the
water height. Each conveys information about the other, but neither
influences the other. If you manipulate, the correlation is gone, and
so is the information you can get about one by measuring the other.

If you DO have a hypothesis that X influences Y, and that the influence is
the reason for an observed correlation, you can check it out by manipulating
X. If you also have a hypothesis that the influence is bidirectional, you
can check that out, too, by manipulating each separately, and seeing if the
correlation still holds. If it does, you may have been right about there
being an influence. If the correlation fails, you were probably wrong. But
when you don't interfere with it, the relationship is still there, even
when there's no mutual influence either way.

I hope this helps a bit.

Martin

[From Bill Powers (970428.0645 MST)]

Martin Taylor 970427 13:30]

Here's another point to mull over. If X provides B bits of information
about Y, Y provides B bits about X. It matters not a whit whether X
influences Y, X and Y are both influenced by Z, Y influences X, or none
of these are true. Information, like correlation, carries no implication
of causality.

Yes, I can accept that. Y provides B bits of information about X, but
this does not mean that by manipulating Y you can induce or predict
changes in X -- unless the function is bidirectional.

No, nor can you cause changes in X by manipulating Y if there's no
influential relationship between them. But the informational relationship
can still hold perfectly well. In the situation analyzed, you may or may
not have been manipulating one or the other. You might have been just
observing. If that was the case, you have changed the system by performing
the manipulation--you've changed the influences, and therefore possibly
the correlation.

Consider this example to see what I mean....

OK, I get the idea. As has been said, correlation does not imply causality,
although causality does imply correlation.

My only hangup is the idea that there is _mutual_ information in a
unidirectional relationship. I can see that while A is influencing B, you
can say that the state of A constitutes information about B (if the moon is
up there will be a tide in 6 hours), and that the state of B constitutes
information about _what A must have been_ (i.e., the high tide indicates
that the moon must have been up 6 hours earlier) -- if nothing else can
influence the water level. Both of these "informations" are really about the
influence of the moon on the tide. They imply nothing about the influence of
the tide on the moon, even though it is said that the tide contains
information about the moon.

I guess I am still wondering how the water level can give information about
the moon if it was raised by pumping. By judicious manipulation of the
pumping rate and direction, you could give the impression that the moon was
anywhere in its orbit, couldn't you?

Well, here I am doing what I said I was going to stop doing; amazing how one
gets sucked into intellectual puzzles even without wanting to. I suspect
that I have some hesitation about actually being able to design some new
experiments that will advance PCT (in the directions I'm interested in) and
am procrastinating.

Dammit, Martin, go solve your puzzle for yourself. I'm just a distraction to
you.

Best,

Bill P.

[Hans Blom, 970428e]

(Bill Powers (970424.0915 MST))

I think that as long as we keep in mind that externally visible
"behavior" has two causes, a fairly straightforward "translation"
between SR-theory and control theory is possible -- although
control theory gives additional insight. It is therefore, maybe,
that PCT impresses many people as "nothing but" something very much
like a top-down versus bottom-up coordinate transformation: useful,
but "nothing new". Alas: they miss the additional insight...

I'm very pleased with your comments here. All is not lost.

What would have been lost otherwise? :wink:

Inside the "disturbance" function in your program, declare a typed
constant (local variable) dd. Make dd := t, so dd increases linearly
with time during the run. Return the value,

disturbance := 10*sin(dd);

I don't get this, Bill. What would dd be, except a copy of t, even if
computed independently? What would it change in the program's output?

Greetings,

Hans

[Martin Taylor 970428 16:40]

Bill Powers (970428.0645 MST)]

OK, I get the idea. As has been said, correlation does not imply causality,
although causality does imply correlation.

No, that's not even true. One point of my example Y = |X| was to suggest
that causality need not imply correlation--at least not linear correlation.

Suppose that this Y happened to be the output of a full wave rectifier
operating on a zero-centred input waveform we label X(t). The "X"s in the
computation of correlation are sample values of X(t), and likewise for
Y, samples from Y(t). Now X _causes_ Y (inasofar as anything can be said
to cause another thing), but X and Y are linearly uncorrelated.

Informationally, though, you are right, so far as I know. Causality
implies positive mutual information between the thing caused and the
causing thing.

Martin

[Frfom Bill Powers (970428.1807 MST)]

Hans Blom, 970428e --

Inside the "disturbance" function in your program, declare a typed
constant (local variable) dd. Make dd := t, so dd increases linearly
with time during the run. Return the value,

disturbance := 10*sin(dd);

I don't get this, Bill. What would dd be, except a copy of t, even if
computed independently? What would it change in the program's output?

I was just trying to get a time-varying disturbance that had a different
form from its effect on the controlled variable. I could have said dd =
t^1.2, or anything else. As it is, we have a disturbing variable that rises
linearly with time, and a disturbing _effect_ that varies sinusoidally.
Obviously, your program reproduces the sinusoid, but it has no way to know
that this effect comes from a disturbance that simply rises linearly.

If I ever get around to it, I think I can show that by calculating the
disturbing _effect_, your program is doing exactly what a PCT model does
without that calculation. The effect of the disturbance gets added into the
output just as it does _implicitly_ in the PCT model. If you have some spare
time, perhaps you could fool with the algebra to see if you can demonstrate
the equivalence.

Best,

Bill P.