[From Mike Acree (2009.02.08.1626 PST)]
In my last exchange with Martin (1/15), I
was asking him for justification for his repeated claim that Bayesianism
represents the ideal of human reasoning, and he was exasperated to the point of
majuscules:
Just what IS it that you want?
I didn’t see what was so impossible
about my question, until I surmised that, from Martin’s point of view,
Jaynes had already provided a proof, so perhaps he thought there was nothing
more to be said. I decided the next step was to read Jaynes and see for
myself, but that took a few weeks.
I thank Martin very much for the
recommendation. Although I was a recovered Bayesian by the time I read his
1976 article, shortly after it was published, I remembered admiring his thought
and enjoying his style; and the present book fully lived up to my expectations.
Jaynes’ practical approach to problems is impressive, and I would be
happy to have him as a consultant on an engineering problem. I am also in
general agreement with his criticisms of the frequentist theory and methods. I
think my only significant disagreement with him lies in the scope of his claims
for Bayesian methods, precisely the point where Martin and I had disagreed.
Jaynes makes the interesting observation
that Bayesians have tended to be physicists, the frequentists mathematicians
and philosophers. The Bayesian physicists would include Harold Jeffreys, his
closest neighbor, but not Keynes or Savage. Jaynes does a very good job of
analyzing physical problems, and identifying relevant symmetries and
invariances for probability assignment; his resolution of Bertrand’s
paradox is a good example. But the broader, foundational claims, of which he
is so proud, are the weak point of the theory.
The basic problem is perhaps nothing more
than confusion of concepts. A minor example is that Jaynes starts off talking
about plausibility. He does this, he says, because he doesn’t want to
bring in at the beginning all the baggage associated with probability. But
careful users of English, of which Jaynes is generally one, distinguish these
two concepts. Whereas probability in everyday language relates to the degree
of support for a proposition, plausibility pertains more to the absence of
evidence for the contrary. That is the way it is explicitly defined in
Shafer’s theory, for example. This conceptual sloppiness reminds me of
the way psychologists, of all people, regularly treat depression and sadness as
synonyms; but in this particular case of Jaynes’ it doesn’t do any
particular harm. The real problem is that Jaynes is constructing a calculus of
truth while thinking of it as a calculus of support. Jaynes, following Cox,
starts from the fact that “A + not-A” is always true, and he
assigns this proposition conventionally a probability of 1. (His derivation of
the additive rule—that probabilities over the various alternatives must
sum to 1—reminded me of how strange it struck me, when I was reading
Cox’s book about 35 years ago, that he would go through several pages of
calculus to derive the sum and product rules.) The problem is that he then
thinks of probabilities in terms of evidence or support. But it is not unusual
to find situations where there is little evidence bearing on a question one way
or the other. Both A and not-A have high plausibility, we might say.
Jaynes’ scale, with certainty at both ends, cannot represent such
situations very well.
Jaynes is smart enough to have recognized
the problem, and his attempt to deal with it is probably the best that could be
done within Bayesian theory. Late in the book, he tell us that he had
considered for awhile the necessity of a two-dimensional logic, encompassing
something like weight of evidence in addition to probability. His own examples
are as useful as any: He would assign a probability of ½ to heads with a coin
he had physically examined and determined to be fair, and also to the
possibility of life on Mars. (The value of ½ in the latter case surprises me,
but it doesn’t matter.) But he acknowledges the difference in the two
situations. In the former he has a large weight of evenly balanced evidence;
in the latter he has (evidently) no evidence at all. He proposes to handle the
situation by giving that probability a distribution, Ap, of which ½ is the mean in each case. For the
coin, Ap would be
sharply peaked, perhaps like a beta distribution with r = 50 and N = 100. For Mars, his Ap
would be U-shaped, like a beta with r
= N = 0. The weight of evidence
is then reflected (inversely) in the variance of the Ap distribution, and will correspondingly be
taken into account in the posterior distribution, which incorporates the
imaginary N generating the Ap distribution. Jaynes is
uncomfortable with the idea of a probability of a probability, and tries to
deal with it by creating an “inner” and “outer” robot,
corresponding, he says, to the subconscious and conscious mind, so that the two
probabilities are on different levels. He is fierce throughout the book in
denouncing the “adhockeries” of frequentism, but this seems an
unfortunate piece of adhockery itself, at the foundations.
Jaynes sees himself from the outset as
developing principles of reasoning for a robot. The robot, in his view, is
what makes his theory logical rather than personal. Jeffreys never won much
acceptance with his logical theory just because it was never clear whose beliefs were being represented.
Jaynes often refers to the robot’s “state of knowledge,” but
the robot as defined is incapable of knowing
anything. It “doesn’t do semantics” (p. 94n), so it
doesn’t know what anything means; that is taken care of by the user.
Jaynes wants the robot to behave like a child and believe everything he tells
it; a skeptical robot would be dangerous (p. 99). The robot will take account
of all evidence and be purely unbiased in its conclusions. All the evidence,
that is to say, as we present it. The most telling example is Jaynes’
discussion of Soal and Bateman’s ESP research. Results with one of their
participants had a p value of 10-137;
like C. E. M. Hansel, Jaynes takes this as proof of cheating. He says
straightforwardly that he would choose whatever prior probability was necessary
to swamp that 10-137; nothing could convince him that ESP was real.
If that isn’t sheer prejudice, I don’t know what is. Perhaps
Jaynes would defend himself by acknowledging that the prejudice resided in him
rather than in the robot; but then it’s not clear what he needs the robot
for—except to make his prejudice look more scientific.
Jaynes thinks of his robot as virtually
human, albeit a slave; he approvingly quotes von Neumann’s challenge:
Specify what a machine can’t do and I will build a machine to do it. Of
course, meeting von Neumann’s challenge is trivially easy (saying,
“Will you marry me?” and meaning
it, for example). But Jaynes’ robot is subject to the usual
limitations, so it is also hard to see from that point of view how Jaynes could
have imagined it an adequate model of reasoning. Any proposition given to the
robot, he says, must have an unambiguous meaning. Setting aside the fact that
meaning is an irrelevant concept to the robot, how many propositions does that
leave—judging just from CSGNet exchanges? Conflicting evidence is also
excluded; we don’t want the robot to undergo the “agony” of
reasoning from incompatible data, which would cause it to crash. But how much
of the real world is then left? Jaynes assures us that evolution selects for
Bayesians, but I think his robot, if it were alive, wouldn’t last long.
Jaynes regularly accuses frequentists of
the “mind projection fallacy,” as in taking randomness to be a
property of things rather than a statement of our lack of knowledge. I’m
prepared to agree with him that the phenomena he is talking about are phenomena
of knowledge; but the problem is that Jaynes’ theory is not epistemic but
alethic. In this regard he could be accused of committing the world projection
fallacy.
Jaynes wants numerical measurement of
plausibilities, or probabilities, so he can program his robot; otherwise he
couldn’t derive all his wonderful results. I’m sure the same point
would be made by modern psychologists, that without numerical measurement of
phenomena like depression or self-esteem, we wouldn’t have been able to
generate the millions of volumes of research filling our libraries. In
Jaynes’ case, I think the useful scientific work could continue, based on
his justifiable arguments about symmetry; all that would be lost are the
unnecessary claims about measuring probabilities in other, ill-defined
situations, like the probability of life on Mars. Jaynes notes that the
desideratum of numerical representation could be replaced by the joint
conditions of transitivity and universal comparability; but Keynes, the first
logical Bayesian, already gave reasons for rejecting the assumption of
universal comparability. Jaynes gives the example of whether it is more likely
that Tokyo will have a severe earthquake on June
1, 2230, or that Norway
will have an unusually good fish catch that day, and gamely argues that they are comparable; but I agree with Keynes
that the general claim is a stretch.
Jaynes boasts frequently that frequentist
techniques are merely special cases of Bayesian theory, so it is odd that he is
so hostile to Glenn Shafer’s theory of belief functions, for which
Bayesian theory is in turn merely a special case. Jaynes does not refer to
Shafer in the text, but describes him in the Bibliography as a “fanatical
anti-Bayesian,” a characterization which would startle Shafer if he didn’t
already know Jaynes. Shafer has the edge on consistency, for his theory
explicitly models belief, or support, or evidence—an epistemic rather
than an alethic approach. It recognizes that, although either A or not-A may
have to be true, we may have, at a given point, little evidence one way or
another, and statements about our knowledge
should accommodate that fact. That said, the Shafer formalism seems to me but
of academic interest. As I indicated earlier, Shafer’s proposals for
numerical measurement of beliefs or support were aptly criticized by Krantz as
analogous to rating the esthetic pleasure derived from viewing a painting by
matching it with the sweetness of a graded series of sucrose solutions of known
concentration.
So I would say, in sum, that Jaynes has
made a terrific contribution to solving practical problems in the physical
sciences, and has provided some provocative and incisive criticisms of
frequentism, but that he has failed in his ambition of providing a general
model of inductive inference.
Mike