Correlation

[From Bruce Abbott (970417.2200 EST)]

Bill Powers (970417.1329 MST) --

Bruce Abbott (970417.1340 EST)

In the section on "Mutual information", Kennaway shows that to predict Y
from X to one part in N (i.e., to distinguish N different values of the
variable), the correlation c must be

c := sqrt(1 - 1/N^2).

To distinguish 2 points (the minimum required to establish a regression
line) the correlation must be 0.866 (for p < 0.05, I presume). But this does
not allow you to assign values on any finer scale. You can say only that if
X is known to be positive, Y will be SOME positive number -- but you can't
say WHAT positive number. That's what "1 bit of Information" means. For 3
points you must have c = 0.94 -- this would enable you to say that if X is
negative, Y is negative, if X is 0, Y is zero, and if X is positive, Y is
positive. To distinguish 5 levels of Y (two negative, two positive, and 0),
you need c = 0.98.

This confuses the accuracy of predicting an observed Y-value given knowledge
of its X-value with the accuracy of pinning down the slope and intercept of
the line used to generate the prediction. Even if we knew the parameters of
the line perfectly, the tendency of points to scatter away from the line
limits one's ability to predict the value of a given observed point's
Y-value from its X-value. The precision of the estimates of the _line's_
paramaters depends directly on the number of independent points in the
sample. The precision of the estimates of a _point's_ position depends on
the precision of estimates of the line's position and the degree of scatter
of Y-values about the X-value in question. The number of bits by which the
uncertainty of a point's value on one variable is reduced by knowledge of
its value on the other has nothing to do with the number of points required
to completely specify a given _function_.

This says to me that drawing a regression line through a scatter plot
representing c = 0.866 is highly misleading; it implies a linear
relationship _between_ the extremes that is unjustified. It takes three
points to establish a quadratic relationship, so you'd need c = 0.94 to get
a _minimal_ indication that Y = X^2 fits the data (you can perform any
transformation you like before applying the "linear" correlation method).

Let us assume that there is a simple function relating Y to X. Let us
further assume that Y is not only affected by X, but also by a small random
disturbance. How would you go about nailing down the true function, given
the ability to sample Y-values at any desired X-value? One way would be to
set X to a given value and then repeatedly sample Y. You would end up with
a normal distribution of Y-values centered about their mean. With a large
enough sample you could estimate the the true value of Y (its value in the
absence of the random disturbance) to any desired degree of precision. You
might then try another X-value and repeat the process, then again, and
again, until you had swept out the function over the range of interest. You
might even be able to find a mathematical formula that perfectly fit the
estimated true Y-values.

Now, computing Pearson r for the set of points you observed in this process,
you discover that it is 0.886. Only one bit of information about the
position of Y given X, and yet you have established that the underlying
function is, say, cubic, with an extremely high level of confidence. Go figure.

Linear models are "powerful" only if you have some independent way of
knowing that the relationship should be linear.

No, I mean by "powerful" the ability to improve prediction relative to what
is possible. In a large number of cases a linear model will work nearly as
well as one embodying the "true" nonlinear function. In otherwords, linear
models often turn out to be excellent approximations, at least over some
practical range of values in which one has an interest.

When you're presented with a
data set and have no knowledge of what the relationship is, the linear model
(where I come from) is considered the _weakest_ assumption, and it's
generally used only because of a lack of justification for anything more
complicated. And also, no less important, because there aren't very many
nonlinear equations that we can solve!

And also because they so often turn out to work surprisingly well even when
they are not actually the correct function. In the absence of better
information, use what is most likely to work (at least tolerably well).

To distinguish an exponential from a straight line, you need at least two
slopes; i.e., three data points, which implies c = 0.94, at least. And that
would be quite insufficient to distinguish the exponential (A*exp(kx)) from
a quadratic (Ax^2 + Bx).

You can't get away from the fact that with a correlation of only 0.866, you
could not distinguish a straight line from an exponential relationship by
ANY means. But I should let Richard have his say on that.

I think you can see now that this is nonsense. But by all means, let
Richard have his say.

You miss the point. Even to see that there are significant deviations from
theory, you must have very good data; the basic correlations must be in the
high 0.9s. To distinguish among different explanations for the deviations,
you must have better data still, or a better model. This means that the
major part of your labor has to go into getting very good data and improving
the model, not into devising ever more "powerful" ways of extracting
information from noisy data.

No, Bill, I don't miss the point. Precise is better than imprecise. I
don't dispute that. But I'm curious: what do you think _my_ point is?

You last sentence has me baffled. The object of observing the >correlations

in the data is to account for the variation in the response variable. You
don't know what "regular function," if any, may be at work. The regular
function is all you need? The regular function is what you seek to >discover!

Bafflement all around, here. It might help if you told us what the trick is!

But there _is no trick!_ What, exactly, do you want me to tell you?

I can show you a nice correlation between Y and sin(X) over the X

range 90 to 270 degrees).

Well, not THAT nice! What is it, by the way? Not that it would mean anything
-- the distribution is not normal.

(I should have said "correlation between X and sin(X); sin(X) is Y.) The
correlation between X (in degrees) and sin(X) is -0.993 using as points only
the integer degrees. The linear relation between X and sin(X) accounts for
98.6% of the variation in sin(X) over the range of X-values used. In other
words, a sawtooth wave is an excellent approximation to a sine wave.
Powerful stuff, these linear fits. And by the way, the correlation _does_
mean something, even though the distribution is not normal.

But whether correlations are low because there are
influential variables not being measured or because the wrong model
(linear) is being applied to the data, those correlations are still
scientifically informative and useful. Furthermore, correlations should
never be interpreted without a look at the scatterplot. If a strong
_non_linear relationship is present, that will be revealed by the plot.

Not with a correlation of only 0.866. Unless you mean a VERY strong
nonlinearity or an obviously non-normal distribution.

With a correlation of 0.866, 75% of the variation in Y is accounted for by
variation in X (or vice versa). In other words, if variation in Y is due to
variation in X, then holding X constant would reduce the observed variation
in Y to 1/4 its former value. Are you saying that a 75% reduction in noise
is valuless? I hope you weren't in the business of designing radio
receivers! Bill to customer: "You don't need those newfangled designs; this
here crystal set is all you need! Why, the newer models only cut the static
by 75% -- hardly worth the price!" (;->

I think that low correlations are generated by bad experiments, so in my
view I didn't change anything. I don't expect people to suddenly start
getting high correlations by improving their bad experiments; I just wish
they would stop PUBLISHING before they have refined their experiments.

Well, I then I guess you and I have been collaborating on bad experiments.
Even wet versus dry fecal weights correlate only on the order of 0.90, and
those are among our best correlations among the measures we collected.

What I _did_ state is that several variables having low correlations with
a response measure sometimes can be combined to yeild a function that
predicts well. As this was _demonstrated empirically_ in the simulation I
reported on in my post, the truth of this statement is beyond question.

I'd still like to see what's behind those numbers. Is this a common
situation? If not, how do you tell when you're dealing with an example of it?

In the simulation I had Minitab generate four samples of
normally-distributed values, which were placed in four columns. Each sample
was generated independently by random sampling. I then summed across the
rows to produce the values of the response variable. Thus the _true_
underlying function was R = P1 + P2 + P3 + P4. I then had Minitab compute
the correlation matrix. Multiple linear regression was then used to fit the
data to the function b0 + b1P1 + b2P2 + b3P3 + b4P4 = R, the coefficients
being estimated from the data. Not surprisingly, Multiple r (the
correlation between predicted and observed R-values) was +1.0, as was
r-squared. Given the way the response variable was derived, this was the
expected result; the point of the simulation was to demonstrate that a full
account of the variation in R could be obtained from variables whose
individual correlations with R were in the range deemed by Kennaway as
"useless." In this case each individual variable could account for only
about 1/4 of the variation in R.

Regards,

Bruce

[From Bruce Abbott (970421.1515 EST)]

Richard Kennaway (970421.1610 BST) --

Bruce Abbott (970419.1740 EST):

Because you are more familiar with concepts like the S/N ratio than I, I
wonder if you could tell me what relationship exists between the S/N and the
proportions of variance accounted and unaccounted for. I have a feeling
that a correlation could be reexpressed as a signal-to-noise ratio, although
information about the sign of the correlation would vanish.

SN power ratio in dB (call it SNdB) is 10*log-base-10(Var R/Var N).

Therefore c-squared = 1/(1 + 1/SN), where SN = 10^(SNdB/10)

Equivalently, SNdB = -10*log-base-10( 1/c^2 - 1 ).

Thanks for the help; it would appear that my intuition was correct! (I
might have done the derivation myself but wasn't clear how signal and noise
are defined in computing the S/N ratio of audio systems (e.g., as variances
vs. rms values).

This way of expressing the meaning of a correlation should be right up Bill
"Mr. Electronics" P.'s alley -- something he can get his teeth into.

   c SNAdB Var R/Var N
   0.2 -13.80 0.042
   0.5 -4.77 0.33
   0.707 0 1
   0.866 4.77 3
   0.9 6.29 4.26
   0.95 9.66 9.26
   0.995 ~20 ~100
   0.9995 ~30 ~1000

Radio engineers are usually working in the region where signal is vastly
greater than noise. Psychologists -- those that work with low
correlations, at least -- work in the region where Var R is of the same
order of magnitude as Var N.

I'm glad you added the qualifier, as psychological research does exist in
which the S/N ratio is much higher, although generally not in the region
radio engineers are used to.

I'll have to defer answering your other posts for the moment as I'm getting
fairly busy.

Regards,

Bruce

[From Bruce Abbott (951002.2135 EST)]

Tom Bourbon (950929.1340) --

For your largest correlation, +.53, the percentage of variance "explained"
is 28%. The coefficient of alienation, k, for that correlation is
SQRT(1 - r-squared), which is .85. That means the degree of lack of
relationship is .85, the degree of relationship, only .53. Another way to
describe this relationship is to say that, when r = .53, the Standard Error
of the estimate for the predicted variable is 85% as great as the margin of
error were you to make the prediction when you knew _nothing_ about the
degree of relationship between the variables. Those are weak data. It is
not your fault they are weak. It is not your fault that behavioral science
is built on a foundation of data that are similarly weak, albeit
statistically significant.

Another way to look at it is that 100 - 28 = 72% of the total observed
variation in one of the variables would remain if the other variable were
held constant (assuming that the relationship between the two variables is
linear). This result suggests that substantial variation in one variable is
unrelated to variation in the other. Let's assume that the two variables
are X and Y and that the relationship is a causal one, with variation in X
producing variation in Y. Obviously, the influence of X on Y is relatively
weak compared to the influence exerted by other variables at work in these
observations. Those other variables have been free to vary across
observations and thus to weaken the apparent relation between X and Y.

Note that the apparent "weakness" of the relationship between X and Y is
only relative: X can have a strong influence on Y, but if other strong
influences are also allowed to act on Y, the observed relationship between X
and Y will appear weak.

If the magnitudes of those uncontrolled sources of influence vary randomly
from one case to another, then despite the weakness apparent in the
relationship it should be possible even with fairly weak correlations to
determine to a reasonable approximation the nature of the underlying linear
relationship between X and Y. The following experiments demonstrate this
point. Minitab was used to generate a normally-distributed error variable
with a population mean of zero and a known population standard deviation.
The X observations were the numbers 1-100; the Y observations were
synthesized by multiplying X by 2 and then adding the random error. Thus, X
and Y were related by the equation

Y = 2X

but this relationship was obscured by the "influence" of uncontrolled error.
I systematically varied the standard deviation of the error, and used linear
regression to estimate the "true" underlying relationship between X and Y.
Here are the results:

Error
Pop. SD Fitted Straight Line Sest R R-sq
   10 Y = +0.64 + 1.98 X 11.16 .982 96.4%
   20 Y = -0.02 + 1.98 X 18.91 .951 90.4%
   40 Y = -3.86 + 2.02 X 38.67 .836 69.8%
   80 Y = -4.90 + 1.88 X 73.59 .598 35.8%
  160 Y = +7.00 + 1.98 X 161.50 .337 11.4%

For making predictions of Y from X, the results get progressively worse as
the amount of variation in error increases relative to the (constant)
variation in X in these experiments. However, if one's intention is to use
the data to infer the "true" relationship between X and Y (assuming that it
is linear), even the worst case shown here provides an excellent estimate of
the true line.

The reason correlations are typically so high in tracking experiments is
that there are no disturbances to cursor position at work during the task
that are anywhere near the size of the disturbance being applied by the
program. But make the cursor "too sensitive" to mouse movements and watch
the numbers deterioriate. It is not the science that makes the correlations
high in the usual tracking experiment, but the absence of any sources of
"noise" potentent enough to seriously disturb the clean relationship between
disturbance and mouse movement. In more complex situations where unmeasured
and uncontrolled (in the experimental sense) sources of error exert as
strong an influence as the disturbance does, control (in the PCT sense) will
be poor and the correlations will be correspondingly less impressive.

Thus, when Tom says,

Whether or not PCT exists,
these correlations are weak and any science built on them must be flawed,
especially if the scientists purport to explain the behavior of individual
people. It is not your fault that most behavioral science is built on
"facts" that are fatally flawed.

I cannot agree entirely. As demonstrated above, it is possible to infer
something about a relationship between two variables even when the
correlation between them is weak, although that inference rests on certain
assumptions which may or may not be reasonable in a given case. I do agree
that one should be looking for relationships within single individuals
rather than across groups of them. Group-based correlations can tell you
what to expect on average from a group of individuals as some variable is
varied, but they provide little definite information about how the
individual responds to those variables. A given group relationship may
emerge because everyone has the same relationship or because everyone has a
different relationship that averages to the group trend.

Regards,

Bruce

[From Bruce Abbott (951003.1135 EST)]

Bill Powers (951003.0500 MDT) --

I think I see one problem in your interesting statistical analysis; it's
one I've run into before, but your analysis made it clearer. When
psychologists talk about the "amount of effect" of one variable on
another, the methods they use always seem to normalize all the variables
to a peak value of 1.00, so any _absolute_ measure of the amount of
effect is lost. Some of this absolute measure is restored by considering
the regression coefficient, but generally there is no way to make sense
of that coefficient.

I'm not sure to what you refer here. If you're talking about r-sq, it does
indeed "normalize" to 1.0 for the simple reason that it is a proportion--the
proportion of variation in one variable that can be "accounted for" by
variation in the other variable. It is a proportionate measure of the
shared variance in the two measures. The correlation coefficient, r,
likewise removes the absolute magnitudes of variation in the two variables:
you will get the same r between X and Y if you add, subtract, multiply, or
divide all scores of either or both variables by a constant; that is, r in
unaffected by linear transformations of the variables. In fact, r is the
regression coefficient for linear regression if X and Y are both transformed
to their normalized (z-score) equivalents.

Inferential tests likewise remove absolute magnitudes from the computations;
thus under the right conditions (e.g., very large samples) a difference
between two means can be statistically significant and yet trivially small.
As with correlation, these tests can be viewed as assessing the proportion
of variation in one variable that can be attributed to variation in another
variable (or variables); it is not the absolute magnitude of an effect that
is relevant in these calculations but the relative magnitude, compared to
the overall variation in the dependent measure. Lately statisticians have
begun to emphasize the importance of examining the data to determine the
absolute size of an effect once it has been determined that the effect is
reliable. This is the important fact, not the smallness of the obtained
p-value.

But there is actually not a great deal of noise in the human system,
only a few percent of the magnitude of the behavior that is affecting
the controlled variable. The reason it seems so large in the statistical
analysis is that it is being compared with the result of adding two very
large variables together in almost perfect opposition -- the disturbance
and the handle position. In absolute terms, the cursor excursions are
only a small fraction of the excursions of the disturbance and the
handle.

Yep.

But when you calculate the effect of the disturbance on the cursor using
the noise-free model, you come up with a rather high correlation because
of the lack of noise in the model. This would say that the disturbance
"accounts for a large part of the variance" of the cursor. This is
perfectly true, but in fact the _magnitude_ of this effect can be made
as small as you please by raising the loop gain of the control system.
So you end up with D having a relatively high correlation with C, at the
same time that D is having only a trivial amount of effect on C.

Correct.

In the control system model we get around this problem by comparing the
variance of the cursor with the variance that would have been expected
if the handle varied randomly relative to the disturbance. This gives us
the "stability factor" that has been mentioned occasionally (see
Spadework article in Psych Rev). However, this actually understates the
degree of control, because if there were actually no control system,
there would be no handle effect to consider and the cursor variations
would be exactly the size of the disturbance variations.

This is a sensible thing to do, given that you know (or can estimate) the
variance to be expected in the absence of control.

What seems to be missing from most psychological treatments is a way of
estimating the size of the effect of A on B in comparison with the
maximum effect that could have occurred.

There are numbers representing "effect size," but these generally provide
something similar to r-sq in that they represent the proportionate size of
an effect relative to the overall variation in the dependent variable. Your
suggestion is an excellent one IF the maximum possible effect could be
determined.

Consider the example in your generated data. You started with X as an
independent variable and Y as a dependent variable. You said that the
underlying relationship was Y = 2*X. But the data you presented was all
normalized: correct me if I'm wrong, but it seems to me that you would
come up with exactly the same numbers if Y = 0.002*X. The standard
deviation, of course is normalized to the range of Y, so the scaling
factor is irrelevant. Yet in the second case, the effect of X on Y is
only 1/1000 of the effect in the first case. In the first line of your
data analysis, X accounts for 96% of the variance in Y in both cases,
but in the first case the effect is about the same size as the
variations in X while in the second the effect is 0.1% as great. So does
X "account for Y" to the same degree in both cases?

Pearson r and r-sq would not change, but the regression equation certainly
would. The numbers that went into the analyses were not normalized: X
varied from 1 to 100 and the error values varied over ranges that depended
on the standard deviation (e.g., the final value, 160, indicates that about
68% of sampled points would fall between + or - 160 (the mean of the error
distribution was zero). The Sest value reported gives the estimated
population standard deviation of the residuals about the line; in each case
it is fairly close to the actual value used to generate the errors. Sest
gives the absolute magnitude of the fluxuation of the points around the
best-fitting line in units of the dependent variable and thus _does_ provide
a non-normalized measure of error.

In a pursuit tracking task, there are two disturbances, the target
movements and the disturbance applied directly to the cursor. Any number
of other disturbances can be applied to the cursor without much
affecting it, as long as the total doesn't get too big or fast-changing.
Even noise from random variations in handle position, introduced by the
person doing the tracking, is reduced by the loop gain of the system
(within its bandwidth of control). The high correlations to which you
refer are between handle position and the sum of all disturbances. What
we look for in detecting control is not a _high_ correlation, but a
_low_ correlation between handle position and cursor position. The low
correlation is seen because people typically keep the error small enough
to be comparable with internal system noise, and it is the system noise
that reduces the correlation of output with input, or disturbance with
input.

If there are disturbances of the cursor other than those we intend
during a tracking experiment, and there always are, they simply add to
the total disturbance. If we estimate the quality of control by
comparing cursor excursions to handle excursions (where only system
noise enters), it doesn't matter how many external disturbances there
are, or whether we have identified them all. The cursor will still vary
far less than we would expect from the handle movements alone, so we can
tell that control is occurring and estimate how good it is (the total
effective disturbance, as Hans Blom has pointed out, can be estimated
from C - H).

Yes, I understand this; apparently I didn't make myself clear. I'm talking
about disturbances the person cannot adequately correct, such as slip-stick
problems in the mouse or arm joints. These would be outside the bandwidth
of the control system. At high mouse sensitivities, such disturbances would
lead to serious over- and under-shoot and reduce the correlation between
handle movement and disturbance.

If there were some external disturbance comparable in size to the ones
we use in the tracking experiments, don't you think it would be rather
easy to find what is causing it?

In many psychological studies the relevant variables may have strong effects
on the observed measure and yet cannot (or in some cases should not) be
adequately controlled. That's why the reliance on statistical methods. If
you want to know whether smoking influences one's chances of getting lung
cancer, you can't control all the potential varibles that might influence
those chances while allowing smoking to vary. However, by measuring those
variables as they vary naturally across people, you can go a long way toward
determining the role of smoking using statistical methods.

I have more to add, but I'm off to Toledo for the day and I've got to get
going. Hans, I had your model in mind when I wrote the previous post on
correlation.

Regards,

Bruce

that question is simply not a question that a control theorist would ask.
that question is IV-DV question. at best it can hint retrospectively
at what might be a controlled variable but i doubt anyone would care to
ask that. again the search for the disturbance--no matter what is the
"effect size"-- in a closed loop situation is equivalent to the search for
the sign stimulus or cues or any other theory that places behavior at the
end of the causal chain. for myself i would ask what is it that cancer has
resulted in. clearly the chemicals inhaled by cigarettes place
certain tissue cultures into a broth that might not be entirely
conducive to the cells of that culture; possibly by interfering with the
enigmatic nutritive demands of the critter by the intruduction of a
chemical that effectively usurps another chemical that is used by the
cell to break down another chemical on which entire metabolic needs
depend (this is actually durn close to the glucose-deprived
situation of e-coli and their subsequent mutuations). then if the
probability of mutuation is a function of the chonic lack of this element
than whoila! cancer and if this result is one that can bring about
change in the depleted stores then we have a closed-loop relation here.

this might be a fairly contrived explanation but
let me say that cancer is a fine example of "control the fact" though my
lightly put "how" pretty much bites. those critters are without a doubt
controlling just great.. put in a noxious situation they have changed
in such a way as to now be surviving in all that crap people put in their
lungs (or factories, or cities, and more IVs!!)and reproducing well also.
unfortunately you might be deep-sixed by their success which is both the
ultimate form of conflict and reveals that they never gave a damn about
you anyway.

i.

···

Bruce Abbott (abbott@CVAX.IPFW.INDIANA.EDU) wrote:

In many psychological studies the relevant variables may have strong effects
on the observed measure and yet cannot (or in some cases should not) be
adequately controlled. That's why the reliance on statistical methods. If
you want to know whether smoking influences one's chances of getting lung
cancer, you can't control all the potential varibles that might influence
those chances while allowing smoking to vary. However, by measuring those
variables as they vary naturally across people, you can go a long way toward
determining the role of smoking using statistical methods.

<[Bill Leach 951005.00:10 U.S. Eastern Time Zone]

Author : kurtzer@UTDALLAS.EDU
Message: 56325 on Wed, 04 Oct 1995 17:36:12 -0500

Bruce:
... chances while allowing smoking to vary. However, by measuring
those variables as they vary naturally across people, you can go a long
way toward determining the role of smoking using statistical methods.

Issac
that question is simply not a question that a control theorist would ask.

That question is simply not a question that a PCTer might not ask. A
control theorist is often concerned with controlling processes that are
not always fully understood.

There is some validity to the idea that "knowing exactly what a specific
electron has done inside a tungstun filament" is not only not useful but
such detailed knowledge my well obscure the "light".

The Alchemy/Chemistry paradigm shift is often used to compare
"Traditional psychology" and PCT (idea originated I believe by Dag and
really is more appropriate than the Potolemy/Copernicus paradigm shift.

A couple of thoughts about that eariler shift might be appropriate here.
In the first place though I don't know how much time was required before
Chemistry became mainstream but I presume that it was a very long time
(of course communications was not very comparable to what exists today
either).

Alchemy brought both success and disaster. It did both before and after
the "introduction" of chemistry.

Much of the "problem" with behavioural issues is IN the paradigm. The
rigorous methods of PCT are NOT as a practical matter applicable to many
of the "important concerns" of people interacting with each other in
their daily lives. Not only that but many of the very real concerns that
we all face even daily may NEVER be answerable in the exacting manner
of PCT.

While it is true that the PCT sense of behavour is that all behaviour is
the behaviour of individuals. Group behaviour is still the individual
behaviour of however many people are in the group. In the final
analysis, so called "group dynamics" is not a function of this mythical
entity the "group" but rather the dynamics of the individuals comprising
this arbitrarily labled "group".

The truth of that however does not exactly deny such things as "on the
average, persons of such and such culture when alledgedly acting to so
and so "common" purpose can be expected to behave in this manner". Is it
accurate? No. Is it "reliable"? Sometimes. Does such help you to
decide if you want to drive to grandma's house 800 miles away on Memorial
Day Weekend as opposed to the previous weekend if you have a choice?

Is it useful to "know" that people in Raleigh "drive differently" (as in
respond differently to traffic conditions and changes in same) than do
people in Los Angeles? Is such "knowledge" applicable to every specific
driver? No, of course not but knowledge about "local customs" is useful.

If you live in a small town (like in say Idaho) that is 95% mormon, I
would suggest that you don't try to make a living with a tobacco and
coffee shop. Again, does that tell you anything about any specific
person (mormon or otherwise)? No, of course not.

Is all such "knowledge" useless? Again, I say no, of course not (but it
is not behavioural science either).

-bill