Interpreting correlations

[From Bruce Abbott (980211.1520 EST)]

Bill Powers (980211.0316 MST) to Dan Miller --

You then opened your file drawer and did the old piece on
why correlations of only .62 can't tell us anything. The problem
with this standard rap is that you are incorrect in your analysis.

Yes, and thanks to Richard Kennaway for giving us the correct numbers. With
a correlation of 0.62, the chances that your explanation will be wrong are
29 percent, as opposed to 50 percent if you were guessing randomly.

Bill, I'd be very interested to know whether Richard Kennaway agrees with
your interpretation here. Richard, what say ye?

Regards,

Bruce

[From Richard Kennaway (980213.1530 GMT)]

Bruce Abbott (980211.1520 EST)

Bill Powers (980211.0316 MST) to Dan Miller --

You then opened your file drawer and did the old piece on
why correlations of only .62 can't tell us anything. The problem
with this standard rap is that you are incorrect in your analysis.

Yes, and thanks to Richard Kennaway for giving us the correct numbers. With
a correlation of 0.62, the chances that your explanation will be wrong are
29 percent, as opposed to 50 percent if you were guessing randomly.

Bill, I'd be very interested to know whether Richard Kennaway agrees with
your interpretation here. Richard, what say ye?

It depends on precisely what Bill means by the phrase "the chances that
your explanation will be wrong".

1. 29% is the chance that the explanation will make the wrong prediction
in an individual case (or, for that matter, when comparing two random
individuals -- a different question but yielding numerically the same
result).

2. The probability that the explanation itself is true is a trickier
notion. It can be subdivided:

2a. The probability that the hypothesised mechanism is present in
everyone, the errors in prediction being due to confounding factors.

2b. The probability that the hypothesised mechanism is present in some
proportion (choose a number, any number) of the population, the errors (and
some of the successes) in prediction being due to confounding factors.

In my opinion there is no useful way to define or measure these
probabilities. To do so requires a lot of additional assumptions and black
Bayesian magic, and any figures you might come up with will say more about
those assumptions than the real world. And I don't see what practical use
they would be. A scientist needs to determine what general statements are
or are not so, not compute probabilities of their truth. "What does 30%
chance of rain mean? Do I carry one-third of an umbrella?"

When testing the existence of a functional relationship with an experiment
which reduces confounding factors to insignificant levels, the hypothesis
is decisively refuted by 29% of wrong predictions.

Going back to the original message that started this, the proposed
explanation was: "the act of reading (and reading a lot) ... creates a
context within which progressive political ideas can generate and thrive"
(Dan Miller (980204.1645)).

I can see no reason to infer this from the 0.62 study. On the contrary,
there is a fundamental error of statistical inference in claiming that a
statistical trend implies that the physical relationship between the two
variables must have the form of a physically real mechanism such as the one
Dan proposes, plus confounding factors. Mathematically, you can draw the
trend line, but physically, there need not be any such thing. Philip
Runkel has written the book on this.

-- Richard Kennaway, jrk@sys.uea.ac.uk, http://www.sys.uea.ac.uk/~jrk/
   School of Information Systems, Univ. of East Anglia, Norwich, U.K.