[From Bruce Abbott (970417.2200 EST)]
Bill Powers (970417.1329 MST) --
Bruce Abbott (970417.1340 EST)
In the section on "Mutual information", Kennaway shows that to predict Y
from X to one part in N (i.e., to distinguish N different values of the
variable), the correlation c must be
c := sqrt(1 - 1/N^2).
To distinguish 2 points (the minimum required to establish a regression
line) the correlation must be 0.866 (for p < 0.05, I presume). But this does
not allow you to assign values on any finer scale. You can say only that if
X is known to be positive, Y will be SOME positive number -- but you can't
say WHAT positive number. That's what "1 bit of Information" means. For 3
points you must have c = 0.94 -- this would enable you to say that if X is
negative, Y is negative, if X is 0, Y is zero, and if X is positive, Y is
positive. To distinguish 5 levels of Y (two negative, two positive, and 0),
you need c = 0.98.
This confuses the accuracy of predicting an observed Y-value given knowledge
of its X-value with the accuracy of pinning down the slope and intercept of
the line used to generate the prediction. Even if we knew the parameters of
the line perfectly, the tendency of points to scatter away from the line
limits one's ability to predict the value of a given observed point's
Y-value from its X-value. The precision of the estimates of the _line's_
paramaters depends directly on the number of independent points in the
sample. The precision of the estimates of a _point's_ position depends on
the precision of estimates of the line's position and the degree of scatter
of Y-values about the X-value in question. The number of bits by which the
uncertainty of a point's value on one variable is reduced by knowledge of
its value on the other has nothing to do with the number of points required
to completely specify a given _function_.
This says to me that drawing a regression line through a scatter plot
representing c = 0.866 is highly misleading; it implies a linear
relationship _between_ the extremes that is unjustified. It takes three
points to establish a quadratic relationship, so you'd need c = 0.94 to get
a _minimal_ indication that Y = X^2 fits the data (you can perform any
transformation you like before applying the "linear" correlation method).
Let us assume that there is a simple function relating Y to X. Let us
further assume that Y is not only affected by X, but also by a small random
disturbance. How would you go about nailing down the true function, given
the ability to sample Y-values at any desired X-value? One way would be to
set X to a given value and then repeatedly sample Y. You would end up with
a normal distribution of Y-values centered about their mean. With a large
enough sample you could estimate the the true value of Y (its value in the
absence of the random disturbance) to any desired degree of precision. You
might then try another X-value and repeat the process, then again, and
again, until you had swept out the function over the range of interest. You
might even be able to find a mathematical formula that perfectly fit the
estimated true Y-values.
Now, computing Pearson r for the set of points you observed in this process,
you discover that it is 0.886. Only one bit of information about the
position of Y given X, and yet you have established that the underlying
function is, say, cubic, with an extremely high level of confidence. Go figure.
Linear models are "powerful" only if you have some independent way of
knowing that the relationship should be linear.
No, I mean by "powerful" the ability to improve prediction relative to what
is possible. In a large number of cases a linear model will work nearly as
well as one embodying the "true" nonlinear function. In otherwords, linear
models often turn out to be excellent approximations, at least over some
practical range of values in which one has an interest.
When you're presented with a
data set and have no knowledge of what the relationship is, the linear model
(where I come from) is considered the _weakest_ assumption, and it's
generally used only because of a lack of justification for anything more
complicated. And also, no less important, because there aren't very many
nonlinear equations that we can solve!
And also because they so often turn out to work surprisingly well even when
they are not actually the correct function. In the absence of better
information, use what is most likely to work (at least tolerably well).
To distinguish an exponential from a straight line, you need at least two
slopes; i.e., three data points, which implies c = 0.94, at least. And that
would be quite insufficient to distinguish the exponential (A*exp(kx)) from
a quadratic (Ax^2 + Bx).
You can't get away from the fact that with a correlation of only 0.866, you
could not distinguish a straight line from an exponential relationship by
ANY means. But I should let Richard have his say on that.
I think you can see now that this is nonsense. But by all means, let
Richard have his say.
You miss the point. Even to see that there are significant deviations from
theory, you must have very good data; the basic correlations must be in the
high 0.9s. To distinguish among different explanations for the deviations,
you must have better data still, or a better model. This means that the
major part of your labor has to go into getting very good data and improving
the model, not into devising ever more "powerful" ways of extracting
information from noisy data.
No, Bill, I don't miss the point. Precise is better than imprecise. I
don't dispute that. But I'm curious: what do you think _my_ point is?
You last sentence has me baffled. The object of observing the >correlations
in the data is to account for the variation in the response variable. You
don't know what "regular function," if any, may be at work. The regular
function is all you need? The regular function is what you seek to >discover!
Bafflement all around, here. It might help if you told us what the trick is!
But there _is no trick!_ What, exactly, do you want me to tell you?
I can show you a nice correlation between Y and sin(X) over the X
range 90 to 270 degrees).
Well, not THAT nice! What is it, by the way? Not that it would mean anything
-- the distribution is not normal.
(I should have said "correlation between X and sin(X); sin(X) is Y.) The
correlation between X (in degrees) and sin(X) is -0.993 using as points only
the integer degrees. The linear relation between X and sin(X) accounts for
98.6% of the variation in sin(X) over the range of X-values used. In other
words, a sawtooth wave is an excellent approximation to a sine wave.
Powerful stuff, these linear fits. And by the way, the correlation _does_
mean something, even though the distribution is not normal.
But whether correlations are low because there are
influential variables not being measured or because the wrong model
(linear) is being applied to the data, those correlations are still
scientifically informative and useful. Furthermore, correlations should
never be interpreted without a look at the scatterplot. If a strong
_non_linear relationship is present, that will be revealed by the plot.
Not with a correlation of only 0.866. Unless you mean a VERY strong
nonlinearity or an obviously non-normal distribution.
With a correlation of 0.866, 75% of the variation in Y is accounted for by
variation in X (or vice versa). In other words, if variation in Y is due to
variation in X, then holding X constant would reduce the observed variation
in Y to 1/4 its former value. Are you saying that a 75% reduction in noise
is valuless? I hope you weren't in the business of designing radio
receivers! Bill to customer: "You don't need those newfangled designs; this
here crystal set is all you need! Why, the newer models only cut the static
by 75% -- hardly worth the price!" (;->
I think that low correlations are generated by bad experiments, so in my
view I didn't change anything. I don't expect people to suddenly start
getting high correlations by improving their bad experiments; I just wish
they would stop PUBLISHING before they have refined their experiments.
Well, I then I guess you and I have been collaborating on bad experiments.
Even wet versus dry fecal weights correlate only on the order of 0.90, and
those are among our best correlations among the measures we collected.
What I _did_ state is that several variables having low correlations with
a response measure sometimes can be combined to yeild a function that
predicts well. As this was _demonstrated empirically_ in the simulation I
reported on in my post, the truth of this statement is beyond question.
I'd still like to see what's behind those numbers. Is this a common
situation? If not, how do you tell when you're dealing with an example of it?
In the simulation I had Minitab generate four samples of
normally-distributed values, which were placed in four columns. Each sample
was generated independently by random sampling. I then summed across the
rows to produce the values of the response variable. Thus the _true_
underlying function was R = P1 + P2 + P3 + P4. I then had Minitab compute
the correlation matrix. Multiple linear regression was then used to fit the
data to the function b0 + b1P1 + b2P2 + b3P3 + b4P4 = R, the coefficients
being estimated from the data. Not surprisingly, Multiple r (the
correlation between predicted and observed R-values) was +1.0, as was
r-squared. Given the way the response variable was derived, this was the
expected result; the point of the simulation was to demonstrate that a full
account of the variation in R could be obtained from variables whose
individual correlations with R were in the range deemed by Kennaway as
"useless." In this case each individual variable could account for only
about 1/4 of the variation in R.
Regards,
Bruce