dashed hopes, correlations

[From Rick Marken (970418.0900 PDT)]

Me:

What do you agree with, Bruce? Martin's notion of how to do
science? Or his remarkable analytic discovery that

there is imperfect information about the disturbance passed by
the perceptual function

Bruce Abbott (970417.1025 EST) --

Answer: I agree with _all_ of Martin's post.

The perceptual function transforms sensory input (s) into a perceptual signal
(p): p = f(s). Assuming the simplest case, where the sensory variable is the
additive result of disturbance and output (as in a tracking task), p = f(o+
d). So I interpret the statement above to mean that the perceptual function,
f(), passes information about d to somewhere (the comparator?). I presume
that this information, once "passed", is carried by some medium (p?) to its
destination. So when you say that there is "imperfect information about the
disturbance passed by the perceptual function" I can't help but think you are
saying that there is information about the state of d in the state of p. But
it is easy to show that there is absolutely no information about d
in p; knowing p tells you nothing about the state of d.

Fortunately, there is no need for the perceptual function to pass information
about d in order for the control system to be able to generate the outputs
that keep p under control. All the control system has to do is continuously
generate outputs, in a closed, negative feedback loop, with the appropriate
amplification and slowing, that are proportional to r-p.

The only point at which I take issue is that Richard's
declaration that low correlations (0.50 or below) are "useless"
may be misunderstood. What he means by that is that they
are useless _individually_ for the _purpose of predicting
the Y-value of a point from its X-value_ (or vice versa).

Now that I've got the paper I see that he is making a much deeper
point as well. Bill alluded to it in his post [Bill Powers (970417.1329
MST)]. Kennaway shows that a correlation less that .86 tells you nothing at
all about the nature of the relationship between X and Y. Since the goal of
most research is to determine the relationship between variables, the
_minimum_ useful correlation is, thus, .86. And this correlation only tells
you the slope of the relationship between two points in the actual
relationship. You have to observe a correlation of at least .98 to tell
whether the actual relationship between X and Y is linear or not.

I think what Kennaway's work shows is that scientists who want to understand
relationships between variables (as we do in PCT) must work to find
correlations greater than, say, .98. Correlations as low as .86 can be useful
in suggesting that there might be _some_ functional relationship between
variables (so further research might be useful). But my understanding of
Kennaway's paper leads me to the conclusion that the only "usefulness" of
correlations less that .86 is that they tell you that you are barking up the
wrong tree.

Undisturbed,

Bruce

That's really the problem (mine, not yours), isn't it? Disturbances like
this have no effect on your controlled variables (such as your perception of
the value of conventional methodology) because you are controlling them so
successfully. People don't reorganize (and become PCTer's, say) unless
disturbances (like the Kennaway paper) DO have a disturbing effect (push the
controlled variable away from its reference).

Bruce Gregory (970418.1035) --

Just to check my understanding: You are saying that the
perceptual signal contains no information about the disturbance
_insofar as the system controlling the perception is concerned_,
are you not?

Correct. All the system itself "knows" about the outside world (d+o)
is p (all the thermostat "knows" about the one aspect of the world it
perceives -- temperature -- is the state of the perceptual representation
of that variable).

(A theromstat can extract no information about disturbances to the
temperature because it does not moniter its own performance. If I
turn the thermostat off, my perception of changing temperature
_does_ provide information about the disturbance, no?)

Precisely. Also, in a hierarchy of control systems (like us) it is possible
for some of those systems to determine which variables are acting as
disturbances to the variables controlled by other systems. There are (high
level) systems in me that can tell that the moving cursor in a tracking task
is a disturbance to the system (in me) that is controlling the distance
between cursor and target. But the system (in me) controlling the distance
between cursor and target knows only the state of that perceptual variable.

Best

Rick

[From Bruce Abbott (970418.1440 EST)]

Rick Marken (970418.0900 PDT) --

About information, Rick sez:

The perceptual function transforms sensory input (s) into a perceptual signal
(p): p = f(s). Assuming the simplest case, where the sensory variable is the
additive result of disturbance and output (as in a tracking task), p = f(o+
d). So I interpret the statement above to mean that the perceptual function,
f(), passes information about d to somewhere (the comparator?). I presume
that this information, once "passed", is carried by some medium (p?) to its
destination. So when you say that there is "imperfect information about the
disturbance passed by the perceptual function" I can't help but think you are
saying that there is information about the state of d in the state of p. But
it is easy to show that there is absolutely no information about d
in p; knowing p tells you nothing about the state of d.

The term "information" in information theory has a quite precise meaning
that would be helpful to keep in mind for this discussion. Let's say that
you have left your umbrella in one of four possible locations (in the front
closet, in your car, in your office at work, or at the restaurant you
visited last night). Each of these possibilities is equally likely, so you
have 2 bits of uncertainty as to the actual location of your umbrella
(because 4 is the second power of 2). You thoroughly search your car and
the closet and come up empty. Your search has provided 1 bit of
"information," because it has cut the uncertainty as to the umbrella's
location in half (there are now only two possible locations left: your
office or the restaurant). "Information" thus conveys the reduction in
uncertainty as to the value of one variable conditional on some event.

If there is "information in the perceptual signal about the disturbance,"
what exactly does this mean? It means that knowing the value of the
perceptual signal would reduce incertainty as to the value of the
disturbance. In other words, there would be some variation in the
perceptual signal that is correlated with variation in the disturbance.

In a perfect control system, this correlation would be zero, because any
effect of the disturbance on the CV would be instantly and perfectly
countered by the system's action. Thus, no information about the
disturbance would be "passed" by the perceptual function. Of course, real
control systems are not perfect, so there will always be _some_ information
about the disturbance available in the perceptual signal. Thus, your
assertion that there is no information about d in p applies only to perfect
control systems, of which nature provides woefully few.

If this perfect system were provided with a signal indicating the current
effect of its own output on the CV, this effect could be subtracted from the
current perceived state of the CV to yield the current perceived effect of
the disturbance on the CV. That is, this new function would reduce
uncertainty about the state of the disturbance to zero, or put another way,
would provide the maximum possible information about the disturbance. (In
real control systems, the effects of lags and other distortions would reduce
the actual information content of this signal with respect to the
disturbance below this optimum.)

Now on to correlation:

The only point at which I take issue is that Richard's
declaration that low correlations (0.50 or below) are "useless"
may be misunderstood. What he means by that is that they
are useless _individually_ for the _purpose of predicting
the Y-value of a point from its X-value_ (or vice versa).

Now that I've got the paper I see that he is making a much deeper
point as well. Bill alluded to it in his post [Bill Powers (970417.1329
MST)]. Kennaway shows that a correlation less that .86 tells you nothing at
all about the nature of the relationship between X and Y.

Bill didn't "allude" to this point; he beat it to death. Did you read my reply?

Kennaway shows nothing of the sort, as he is talking specifically about the
ability to predict a _single_ observed value of Y given the point's X value.
This is an entirely separate issue from being able to determine the _nature
of the relationship_ between X and Y. I took the trouble to provide a
detailed example showing how one could in theory determine the _nature of
the relationship_ between X and Y to _any desired degree of accuracy_ while
simultaneously being unable to say much about the actual value of a _given_
point's Y value knowing its X value. But I suppose that little facts like
these, being inconvenient to your position, are best ignored.

Since the goal of
most research is to determine the relationship between variables, the
_minimum_ useful correlation is, thus, .86. And this correlation only tells
you the slope of the relationship between two points in the actual
relationship. You have to observe a correlation of at least .98 to tell
whether the actual relationship between X and Y is linear or not.

You've _got_ to be kidding! Where did you say you went to grad school?

I think what Kennaway's work shows is that scientists who want to understand
relationships between variables (as we do in PCT) must work to find
correlations greater than, say, .98. Correlations as low as .86 can be useful
in suggesting that there might be _some_ functional relationship between
variables (so further research might be useful). But my understanding of
Kennaway's paper leads me to the conclusion that the only "usefulness" of
correlations less that .86 is that they tell you that you are barking up the
wrong tree.

Wow. I wonder who _shares_ this opinion. Besides Bruce Gregory, that is! (;->

Regards,

Bruce

[From Bruce Gregory (970418.1620 EST)]

Bruce Abbott (970418.1440 EST)

>Rick Marken (970418.0900 PDT) --

>I think what Kennaway's work shows is that scientists who want to understand
>relationships between variables (as we do in PCT) must work to find
>correlations greater than, say, .98. Correlations as low as .86 can be useful
>in suggesting that there might be _some_ functional relationship between
>variables (so further research might be useful). But my understanding of
>Kennaway's paper leads me to the conclusion that the only "usefulness" of
>correlations less that .86 is that they tell you that you are barking up the
>wrong tree.

Wow. I wonder who _shares_ this opinion. Besides Bruce Gregory, that is! (;->

All I know is what I read in _Casting Nets and Testing
Specimens_.

Bruce

[From Rick Marken (970418.1600)]

Bruce Abbott (970418.1440 EST)

The term "information" in information theory has a quite precise
meaning

Yes. I know.

If there is "information in the perceptual signal about the
disturbance," what exactly does this mean? It means that
knowing the value of the perceptual signal would reduce
incertainty as to the value of the disturbance.

I'm with you.

In other words, there would be some variation in the
perceptual signal that is correlated with variation in the
disturbance.

Even when there is an observed correlation between d and p it is
still not correct to say that p reduces the system's uncertainty
about d. The correlation between d and p depends on many factors
(undetectable by the system itself) that are changing over time
(feedback function, disturbance function, loop gain, etc). This means
that the conditional probability of a particular value of d given a
particular value of p (P(d|p)) is changing all the time. The
informativeness of p is proportional to the distribution of P(d|p) over
p. But if this distribution is changing all the time I can't see how it
is possible to talk about p being informative. The probability of a
particular d given a particular p is unknown at the time p is observed
because P(d|p) is always changing.

Thus, your assertion that there is no information about d in p
applies only to perfect control systems, of which nature provides
woefully few.

No. It applies to all closed loop systems.

Me:

Kennaway shows that a correlation less that .86 tells you nothing at
all about the nature of the relationship between X and Y.

Bruce:

Bill didn't "allude" to this point; he beat it to death. Did you
read my reply?

Yes. I still think you are missing the point. But, hopefully, Richard
Kennaway can straighten me out on this when he decides to join the fray.

I took the trouble to provide a detailed example showing how one
could in theory determine the _nature of the relationship_ between
X and Y to _any desired degree of accuracy_ while simultaneously
being unable to say much about the actual value of a _given_
point's Y value knowing its X value. But I suppose that little
facts like these, being inconvenient to your position, are best
ignored.

Could be. My position is that low correlations (less than .98)
mean that you should try to get better data. I would be surprised
if you did show what you say you showed (in which post did you do this?
how can I know the nature of the relationship between variables and not
be able to predict the value of one variable from the other quite
accurately?) but it would not be relevant to my position
anyway -- unless you showed that you could determine the actual
nature of the relationship between X and Y when the correlation between
these variables is less than .86.

Me:

You have to observe a correlation of at least .98 to tell whether
the actual relationship between X and Y is linear or not.

Bruce:

You've _got_ to be kidding! Where did you say you went to grad
school?

I got my degree from the free university of PCT;-)

Me:

But my understanding of Kennaway's paper leads me to the
conclusion that the only "usefulness" of correlations less
that .86 is that they tell you that you are barking up the
wrong tree.

Bruce:

Wow. I wonder who _shares_ this opinion. Besides Bruce
Gregory, that is! (;->

There might be six of us now. I was hoping beyond hope that you
would make it The Magnificent Seven, but I'm much better now;-)

Best

Rick

[From Bruce Abbott (970418.2005 EST)]

Rick Marken (970418.1600) --

Bruce Abbott (970418.1440 EST)

In other words, there would be some variation in the
perceptual signal that is correlated with variation in the
disturbance.

Even when there is an observed correlation between d and p it is
still not correct to say that p reduces the system's uncertainty
about d. The correlation between d and p depends on many factors
(undetectable by the system itself) that are changing over time
(feedback function, disturbance function, loop gain, etc). This means
that the conditional probability of a particular value of d given a
particular value of p (P(d|p)) is changing all the time. The
informativeness of p is proportional to the distribution of P(d|p) over
p. But if this distribution is changing all the time I can't see how it
is possible to talk about p being informative.

I rather doubt that all these changes are generally severe enough to prevent
_any_ variation in disturbance from being reflected in at least some small
variation in p. But there is no use arguing about it; it is an empirical
question. We could answer it with appropriate data from a simple tracking
study.

I note that you had _no_ comment about my claim that a system provided a
sensor to detect the effect of its own actions on its perception of the CV
would be capable of providing almost full information about the disturbance.
Don't want to admit I have it right?

Bill didn't "allude" to this point; he beat it to death. Did you
read my reply?

Yes. I still think you are missing the point. But, hopefully, Richard
Kennaway can straighten me out on this when he decides to join the fray.

Speaking of whom, we haven't heard much from him lately. I hope I haven't
scared him off . . . Maybe he just doesn't want to embarrass you in public.
(;->

I took the trouble to provide a detailed example showing how one
could in theory determine the _nature of the relationship_ between
X and Y to _any desired degree of accuracy_ while simultaneously
being unable to say much about the actual value of a _given_
point's Y value knowing its X value. But I suppose that little
facts like these, being inconvenient to your position, are best
ignored.

Could be. My position is that low correlations (less than .98)
mean that you should try to get better data. I would be surprised
if you did show what you say you showed (in which post did you do this?
how can I know the nature of the relationship between variables and not
be able to predict the value of one variable from the other quite
accurately?) but it would not be relevant to my position
anyway -- unless you showed that you could determine the actual
nature of the relationship between X and Y when the correlation between
these variables is less than .86.

It was just a verbal description -- a thought experiment if you will. The
result seems so obvious I didn't think it required any further proof. But I
can demonstrate it empirically if you wish.

Where did you say you went to grad school?

I got my degree from the free university of PCT;-)

Well at least it didn't cost you anything (being free). No doubt the text
for your statistics course was _Casting Nets_. (It's the only one on the
approved list.)

Wow. I wonder who _shares_ this opinion. Besides Bruce
Gregory, that is! (;->

There might be six of us now. I was hoping beyond hope that you
would make it The Magnificent Seven, but I'm much better now;-)

I'm still hoping for a better position as one of the Four Horsemen of the
Apocalypse. Just who _are_ your Six Pundants of PCT?

Regards,

Bruce

[Martin Taylor 970419 16:00]

Bruce Abbott (970418.1440 EST)] to Rick Marken (970418.0900 PDT) --

Now that I've got the paper I see that he is making a much deeper
point as well. Bill alluded to it in his post [Bill Powers (970417.1329
MST)]. Kennaway shows that a correlation less that .86 tells you nothing at
all about the nature of the relationship between X and Y.

Bill didn't "allude" to this point; he beat it to death. Did you read my
reply?

Kennaway shows nothing of the sort, as he is talking specifically about the
ability to predict a _single_ observed value of Y given the point's X value.
This is an entirely separate issue from being able to determine the _nature
of the relationship_ between X and Y. I took the trouble to provide a
detailed example showing how one could in theory determine the _nature of
the relationship_ between X and Y to _any desired degree of accuracy_ while
simultaneously being unable to say much about the actual value of a _given_
point's Y value knowing its X value. But I suppose that little facts like
these, being inconvenient to your position, are best ignored.

Right. Butting in again, I might point out that if one measurement of X
gives you one bit of information about Y (reduces its standard deviation
by half), then two measures of X give two bits, and ten measures give
ten bits. If you measure, say, a couple of dozen different values of X
scattered around, you've got lots of information to distinguish between
linear, quadratic, exponential or other curves, provided that the curve
parameters can be distinguished with the increased precision your couple
of dozen bits provide you. Correlations of 0.86 give you one bit per
measure, not one bit per experiment.

Martin

[From Bill Powers (970319.2330 mst)]

Martin Taylor 970419 16:00 --

I might point out that if one measurement of X
gives you one bit of information about Y (reduces its standard deviation
by half), then two measures of X give two bits, and ten measures give
ten bits. If you measure, say, a couple of dozen different values of X
scattered around, you've got lots of information to distinguish between
linear, quadratic, exponential or other curves, provided that the curve
parameters can be distinguished with the increased precision your couple
of dozen bits provide you. Correlations of 0.86 give you one bit per
measure, not one bit per experiment.

You seem to be saying that if you make just 10 measurements of Y given X,
you will know the relationship between X and Y to one part in 1024 for any
X. This is for an underlying correlation of 0.866 between X and Y -- as
determined from many measurements of X and Y.

My understanding is that with a correlation of 0.866, the noise is about 1/3
of the signal, RMS (75% of variance accounted for, 25% unaccounted for).
Making 10 measurements using a single value of X should, if I remember
right, reduce the effective noise in the corresponding Y by about sqrt(10),
or about a factor of 3, so the noise should then be about 1/9 of the signal
-- not 1/1024 (2^-10) of the signal. Is my understanding wrong?

Maybe what you're saying is that if you want to know whether the data are
fit the best by a straight line, a quadratic, a cubic, and so on, this
choice can be made with a high degree of discrimination given relatively few
points. However, I think this is a different question from asking how
accurately one can predict a single Y given a single X. What Richard finds
is that given X, one can successfully predict only the _sign_ of Y in a
single case. I don't think you've disagreed with that.

A correlation constructed by comparing X against the _mean_ of a set of
measurements of each Y will yield a higher correlation than just using the
raw data, won't it? I don't know, I'm asking.

I understand that a correlation of 0.866 is rather high by normal standards,
but it seems to me that you're getting more out of this kind of data than is
possible.

While we're here, I'd like to raise again my argument about the uselessness
of low correlations. Here the "uselessness" is specifically in relation to
constructing logical deductions from statistical results. In any deduction
in which several statements have to hold true at the same time, the
truth-value of the conclusion is the product of the truth-values of the
individual statements. I don't need to repeat the whole argument; the point
is that in order for deductions to be likely to be correct, the component
statements must have a high probability of being true. The break-even point
(where flipping a coin would give results as good as those of a logical
deduction) is simply the N-th root of 0.5, where N is the number of
statements that must simultaneously be true. If there are 4 statements, each
must have a truth-probability of 0.84, and so on.

Maybe Richard would oblige by calculating the relationship between
correlation and probability of truth of a statement "Y = f(X)" -- if that
makes any sense. Or tell us what way of putting it does make sense.

Best,

Bill P.

[Martin Taylor 970420 23:57]

Bill Powers (970319.2330 mst)]

Martin Taylor 970419 16:00 --

I might point out that if one measurement of X
gives you one bit of information about Y (reduces its standard deviation
by half), then two measures of X give two bits, and ten measures give
ten bits. If you measure, say, a couple of dozen different values of X
scattered around, you've got lots of information to distinguish between
linear, quadratic, exponential or other curves, provided that the curve
parameters can be distinguished with the increased precision your couple
of dozen bits provide you. Correlations of 0.86 give you one bit per
measure, not one bit per experiment.

You seem to be saying that if you make just 10 measurements of Y given X,
you will know the relationship between X and Y to one part in 1024 for any
X. This is for an underlying correlation of 0.866 between X and Y -- as
determined from many measurements of X and Y.

I can't believe I wrote it like that! But I did, and have to answer for it.

My understanding is that with a correlation of 0.866, the noise is about 1/3
of the signal, RMS (75% of variance accounted for, 25% unaccounted for).
Making 10 measurements using a single value of X should, if I remember
right, reduce the effective noise in the corresponding Y by about sqrt(10),
or about a factor of 3, so the noise should then be about 1/9 of the signal
-- not 1/1024 (2^-10) of the signal. Is my understanding wrong?

No, I don't think you are wrong. But what has to be added to my original
statement is that you get one bit per measure if the correlation with
_what is still unpredicted_ remains 0.866. However, with repeated measures,
the correlation is with the original data, not with the unpredicted
residual. That means that the second measure is in part predictable
from the first, and does not contribute one more bit to the measurement.
Another way of looking at it is to see that the second measure's value
is in part predictable from the result of the first measurement. It can't
contribute as much informaion to the precision of the Y value as the
first measurement did.

By not making this clear at the outset, I obviously caused you to miss
the point of the posting, which is that one bit of information in a
continuous measurement cannot be equated to saying that you can only get
the sign right. You can't even get the sign right if the distributions are
Gaussian. There will always be cases in which the actual sign is opposite
to what you judge it to be.

Rather than enabling you to judge the sign of a measurement, one bit
allows you to reduce the standard deviation of your probability density
function by half.

···

----------------

While we're here, I'd like to raise again my argument about the uselessness
of low correlations. Here the "uselessness" is specifically in relation to
constructing logical deductions from statistical results. In any deduction
in which several statements have to hold true at the same time, the
truth-value of the conclusion is the product of the truth-values of the
individual statements. I don't need to repeat the whole argument; the point
is that in order for deductions to be likely to be correct, the component
statements must have a high probability of being true. The break-even point
(where flipping a coin would give results as good as those of a logical
deduction) is simply the N-th root of 0.5, where N is the number of
statements that must simultaneously be true. If there are 4 statements, each
must have a truth-probability of 0.84, and so on.

This isn't wrong, but I think it's wrongly applied. I've thought so every
time you bring it up. The problem is that we usually aren't dealing in
logical functions, in which a proposition is false if it isn't true.
Usually we are dealing in approximations that are better or worse. And
in those circumstances, gaining information is always valuable. The only
real question is whether it is cost-effective. All the arguments about it
being better to go at a problem from a different viewpoint are arguments
about cost-effectiveness. You get a lot more out of an approximate but
different view than you get out of a lot of refinement of the old view.

Go back to my original mis-statement. If Y is correlated 0.866 with X,
and also with Z (Z being uncorrelated with X), then you should expect
to get one bit of information about Y from measuring X, and another
from measuring Z (once each).

Martin

[From Richard Kennaway (970421.1600 BST)]

Martin Taylor 970419 16:00:

Right. Butting in again, I might point out that if one measurement of X
gives you one bit of information about Y (reduces its standard deviation
by half), then two measures of X give two bits, and ten measures give
ten bits.

There seems to be some confusion here. Ten measures of X give you ten bits
of information about the corresponding ten values of Y, one bit each. If Y
= Y1 + Y2, where Y1 is some function of X and Y2 is noise unrelated to X,
then most of those ten bits is not information about the function Y1.
Besides, I'm not sure what "information about a function" means,
mathematically speaking, unless one has some prior probability distribution
over the space of functions. If Y1 is, say, a cubic polynomial whose
unknown coefficients are each drawn from some known distribution, and Y2 is
Gaussian noise with a similar magnitude to Y1, then you will need MUCH more
than 10 samples to get 10 bits of information about the coefficients. But
I can't think of a realistic example where you would know the distributions
of the coefficients a priori.

···

__
\/__ Richard Kennaway, jrk@sys.uea.ac.uk, http://www.sys.uea.ac.uk/~jrk/
  \/ School of Information Systems, Univ. of East Anglia, Norwich, U.K.

[From Richard Kennaway (970421.1615 BST)]

Bill Powers (970319.2330 mst):

Maybe Richard would oblige by calculating the relationship between
correlation and probability of truth of a statement "Y = f(X)" -- if that
makes any sense. Or tell us what way of putting it does make sense.

I'm not sure what you have in mind there. But recalling my posting
[Richard Kennaway (970421.1615 BST)], suppose we know (somehow) that Y =
f(A,X) + Y', where Y' is Gaussian noise independent of X, f is a known
function, and A is an unknown parameter to that function with a known
probability distribution. We are given some observations of values of X
and the corresponding observed values of Y, all for a fixed but unknown
value of A. We wish to estimate A from these observations. How much
information do the observations give us about A?

I don't have an answer to this off the top of my head, but it looks like a
meaningful question. Is it anything like the question you were trying to
ask?

···

__
\/__ Richard Kennaway, jrk@sys.uea.ac.uk, http://www.sys.uea.ac.uk/~jrk/
  \/ School of Information Systems, Univ. of East Anglia, Norwich, U.K.

[From Bill Powers (970421.1348 MST)]

Richard Kennaway (970421.1615 BST)--

Keep the comments coming; they are very illuminating.

Maybe Richard would oblige by calculating the relationship between
correlation and probability of truth of a statement "Y = f(X)" -- if that
makes any sense. Or tell us what way of putting it does make sense.

I'm not sure what you have in mind there. But recalling my posting
[Richard Kennaway (970421.1615 BST)], suppose we know (somehow) that Y =
f(A,X) + Y', where Y' is Gaussian noise independent of X, f is a known
function, and A is an unknown parameter to that function with a known
probability distribution. We are given some observations of values of X
and the corresponding observed values of Y, all for a fixed but unknown
value of A. We wish to estimate A from these observations. How much
information do the observations give us about A?

I don't have an answer to this off the top of my head, but it looks like a
meaningful question. Is it anything like the question you were trying to
ask?

This may be what I meant, but I don't recognize it. Let me try again.

Basically, I'm asking about categorical predictions. The hypothesis is, say,
that people with incomes over $50K per year will tend to be Republicans -
i.e., Tories. There is a correlation of c between admitting to an income of
over $50 K and admitting to be a Republican. What are the chances that the
next person you meet who has an income of over $50K will turn out to be a
Republican? Or in other words, we have this syllogism:

People who actually have incomes of over $50K are actually (that is, vote
as) Republicans.

John says he has an income of over $50K.

Therefore John will say he is a Republican.

My question is, given all the correlations involved in the major and minor
premises, what is the probability that the conclusion is true? The
underlying question is, how high do the correlations have to be to permit
the drawing of reasonably valid conclusions when a number of conditions have
to hold true at once?

Note that there are some hidden correlations:

John votes as a Republican vs John says he is a Republican

John actually has an income over $50K vs John says he has an income over $50K.

I contend that

(a) Correlations are generally lower than reported, because the uncertainty
in the measures of the variables is greater than zero, and

(b) Conclusions based on statistical data are less reliable than they appear
from the naive application of logic.

Best,

Bill P.