Martin's post on info theory

[From Bill Powers (960703.1930 MDT)]

Martin Taylor 960703 10:30 --

Your exposition on probability is coming along nicely. Very clear, and
simple enough for me to follow -- so far. I appreciate the fact that
you're open about the basic subjectivity of the concept. I think this
will go a long way toward helping to distinguish those situations where
a probabilistic treatment of phenomena is appropriate, and where it is
not.

You won't see this until you get back, but I'll record some notes
anyway, just while I have your post in front of me.

···

---------------------------
The concept of probability seems to rest on the idea that in identical-
seeming circumstances, a process can have more than one outcome. I think
this is a folk idea, however, that is subject to some doubt. You say

     We say that one toss of a given coin is like another of the same
     coin, and that a result of heads is as probable as a result of
     tails. We say this because we have no reason to believe that any of
     the factors that matter to whether the coin lands on one side or
     the other have changed between tosses.

I don't think this is quite the situation. If all the factors affecting
the fall of the coin were the same, it would end up the same way every
time (almost). A supercomputer simulation given precise starting
conditions could probably predict the coin's final position with
considerable success. Most tosses would involve landing conditions well
away from the edge-on case, so the bifurcation of causality we would
expect for that case would have no material effect on the prediction.

I think what probability means is an inability to predict an outcome,
because of a combination of imprecision in creating or measuring a known
starting condition, and a degree of complexity in the prediction itself
that is beyond normal human capacities to handle. If I flipped a coin so
it executed a single half-turn before landing, that would not be
considered a fair toss. What we expect is for the coin to spin so many
times, and bounce around enough after landing, that nobody could
manipulate the flip or compute the outcome accurately enough to predict
the result. In that case, we attribute the outcome to "chance," but of
course all we mean is that any prediction we might make would be wrong
as often as right. Our prediction would be based on factors that are
unrelated to the factors that determine the coin's final position. You
hint at the same idea:

     Whatever our perception of the probability is based on, it is a
     personal perception, no more veridical with respect to a property
     of the factual world than is any other perception.

It is the lack of any relationship between the personal perception and
the actual determining factors that creates the notion of "chance." What
makes a gambler (amateur) is the false impression that the gambler _can_
predict an outcome; that there is _something_ the gambler can do that
will make the gambler's prediction come true, such as wearing the lucky
hat or picking the horse's number from a list of birth dates. When the
gambler fails to predict correctly, the blame is not placed on his
ignorance; it is placed on an implacable force called "chance" which
says that even if you do everything exactly the same every time,
"chance" will cause differences in the outcome. The illusion, of course,
is in the conviction that you have done everything exactly the same
every time.

Of course some gamblers don't believe in chance. They study frequencies
of occurrance and adjust their predictions so they become businessmen,
not gamblers. They don't care _why_ outcomes vary; they simply study
_how_ they vary.
------------------------------
You bring up a very important point with this:

     To assess a probability, the range of possible results of an event
     must be determined.

I think this is one key to determining when probabilities can be used
and when they can't. If you have a set of mutually-exclusive
possibilities _that you can completely enumerate_, then the rest of the
body of probabilistic calculations follows. The hitch, of course, is in
this complete enumeration. Generally this is possible only when the
world has been reduced to a finite model, so there can be no unforseen
outcomes. If this model is adequate, all the possibilities that matter
can be listed, and valid simple and conditional probabilities can be
observed or generated theoretically ("_a priori_") from the model.

The biggest objections I find to probabilistic treatments comes from the
use of a model that is obviously simpler than the observed processes it
is supposed to represent. It's like saying "Look, either the Moon is
made of green cheese or it isn't, so there's a 50% chance that the Moon
is made of green cheese." A too-simple model misrepresents the
distribution of probabilities, presenting a truncated list of
possibilities as the totality. When the list is not complete, none of
the calculations of simple or conditional probabilities means anything.

The other key to determining when probabilities can be used is to make
sure that the list of frequencies of occurrance can't be derived from
any deterministic model. If you treat a sine-wave as a random variable,
you can record the frequency with which the value will be within some
small range of a particular value between 1 and -1. The only problem is
that the value of the sine-wave is predetermined by its initial value,
and all subsequent values are exactly predictable. Among other things,
variations in the value of the sine wave will not add in quadrature to
the variations in any true random variable -- none of the probabilistic
manipulations applied to relations among random variables will be valid.
This is why our use of "correlations" in characterizing tracking
experiments is, strictly speaking, incorrect. A major part of the
modeled and observed variations can be accounted for by a systematic
model. Only deviations of predictions from the model can properly be
treated statistically.
-------------------------------
I think your argument begins to weaken when you try to establish a
relationship between subjectively estimated probabilities and
mathematically defined probabilities. As I pointed out above, the
essence of the concept of chance is our inability to predict or control
an outcome, but our refusal to blame that on our own ignorance. You say,

     Another basic statement about probability is a mathematical one.
     The subjective feeling of probability ranges from being essentially
     certain that the event will happen to being essentially certain
     that it will not. Mathematics can't deal very well with
     "essentially certain" but it deals well with numbers. So we assign
     numbers with conventionalized meanings. Zero probability means
     certainty that the event will not have the designated result, and
     unity means certainty that it will.

In mathematics, zero probability doesn't mean _a feeling of certainty_
that the event will not happen. It means that the event is treated as if
in fact it never happens. But subjective feelings of certainty most
often do not jibe with observations of frequencies of occurrance. I'm
sure you're more familiar than I am with all the studies in which
subjective estimates of probabilities are found to be simply invalid,
often grossly missing the mathematically-calculated probabilities. The
factors on which subjective feelings of certainty are based obviously
include far more variables than the mathematical calculation of
uncertainty includes. The amateur poker player draws to a straight flush
because he figures he's about due for a bit of luck, and so on. Even the
subjective _meaning_ of certainty is different from the mathematical
meaning; people often whip up a sense of certainty about an outcome
simply to help them achieve it: "I KNOW I am going to win this race!"

Essentially the same comments apply to subjective versus mathematical
concepts of uncertainty.
-------------------------------
I think it would be interesting to try substituting some other word for
"information." For example, if we simply used the symbol A for it, how
many people would guess that everything you say about A applies to what
people get from hearing and reading other people's words? You might say
that if you are told a telephone number twice, the second repetition
does not increase the measure of A. Most people would probably shrug and
accept it. But if you told them that the second repetition gave them no
added "information" about the telephone number (for example,
012033727213, and a week later, 012033727213), they would wonder what
sort of mastermind you are claiming to be.
-------------------------------
Well, that's all I can sustain for tonight. Mary and I spent five hours
this morning hiking a very steep mountain trail and biting off more than
we could chew. 1.2 miles horizontally, 1100 feet vertically. I guess you
have to do this twice a week to avoid being brought low. Beautiful
scenery, though.
-----------------------------------------------------------------------
Bruce Abbott: got the latest data, will do something with them tomorrow.
-----------------------------------------------------------------------
Best to all,

Bill P.

[Hans Blom, 960703]

(Bill Powers (960703.1930 MDT)) to (Martin Taylor 960703 10:30)

This:

The biggest objections I find to probabilistic treatments comes from the
use of a model that is obviously simpler than the observed processes it
is supposed to represent. It's like saying "Look, either the Moon is
made of green cheese or it isn't, so there's a 50% chance that the Moon
is made of green cheese." A too-simple model misrepresents the
distribution of probabilities, presenting a truncated list of
possibilities as the totality. When the list is not complete, none of
the calculations of simple or conditional probabilities means anything.

And this:

This is why our use of "correlations" in characterizing tracking
experiments is, strictly speaking, incorrect.

You recognize the problem, yet it has created some pretty pernicious
tangles. For instance, how can a discussion about control quality
ever succeed if you keep repeating "but we already have a correlation
of 99%; what is an improvement to 99.5% worth?"

Do you see the relation?

Greetings,

Hans

[From Bill Powers (960704.0700 MDT)]

Hans Blom, 960703 --

     For instance, how can a discussion about control quality ever
     succeed if you keep repeating "but we already have a correlation of
     99%; what is an improvement to 99.5% worth?" Do you see the
     relation?

Using correlations was only intended to communicate with people who
treat all hyopothesis-testing in terms of correlations and related
measures. The most common way we PCT modelers speak among ourselves is
in terms of RMS error of prediction. The model fits (and predicts) the
actions of the person within 5% RMS, or 3%, or whatever the number is.
When you go mechanically through the correlation calculations, these
numbers correspond to correlations of 0.98 to 0.995.

As our tracking data now stand, the best model matches the real behavior
over a modest range of conditions within about 3% RMS. To achieve that
level of fit, we use a model with an integrating output, and a transport
lag. Without the transport lag the prediction error is about 5% RMS;
with it, the error drops to 3%. The remaining differences between the
model handle traces and the real ones don't seem to have any systematic
components, although it's clear that the model behavior is smoother than
the real behavior. There is still some room for improving the model.
Perhaps your suggestion of adding some proportional component would
reduce the matching error further. The problem is that as more
parameters are introduced, the same data are being used to make more
adjustments, and we become more vulnerable to the "curve-fitting"
accusation (given enough parameters, you can fit any theoretical curve
to any data). To achieve a prediction error of 3 to 5 per cent with only
two adjustable parameters is, I think, pretty impressive.

The model, of course, can control much better than the person can, given
the right adjustments of parameters. So we know that we're not just
trying to match the model to some perfect control system and falling
short by 3 to 5 per cent.

It would be interesting to explore wider ranges of conditions and look
into ways to model adaptation. But that sort of work is for a future
time when there are people doing this work with institutional support.
The "modern control theory" approach isn't, to my knowledge, concerned
with modeling actual human behavior, so the particular form of the model
that you have described hasn't been evaluated as a model of human
control processes, adaptive or otherwise. Am I wrong about that? If so,
I'd like to see the data.

···

-----------------------------------------------------------------------
Best,

Bill P.

[Hans Blom, 960705d]

(Bill Powers (960703.1930 MDT)) to (Martin Taylor 960703 10:30)

Martin being absent, let me try to answer some of your points. Not
that I can replace Martin, but to continue the discussion even while
he is absent.

The concept of probability seems to rest on the idea that in identical-
seeming circumstances, a process can have more than one outcome. I think
this is a folk idea, however, that is subject to some doubt.

Even more basic is the notion "identical" itself. Although ancient
wisdom says that you cannot step into the same river twice, yet we
have the urge to see similarities and invent classes or categories.
These are utterly subjective: what may be the same for one person
may be extremely different for someone else. One might say that this
has to do with the perceptual input functions that we have developed.
In another sense, we are frequently helped by discovering
commonalities (going up a level of abstraction, one could say). In
mathematics, for instance, one can talk about the "class of
functions" with the common property that all functions perform a
mapping, although the individual mappings may be very different. Yet,
operations may be discovered that apply to _all_ functions. All words
do the same, basically. The word "human" indicates a class, so that
we can talk about properties common to humans. So does "chair",
"love", and "control". A basic property of many of those classes is
that they cannot be crisply delimited. The reason is, of course, that
"identical" is only an abstract notion, a model (simplification)
itself.

In one sense, no experiment can ever be reproduced. In another sense,
we talk about "repeating" all the time. And that is useful; only
where we assume invariance can we perform some sort of "averaging"
and obtain "knowledge" that can provide some sort of prediction of
how "identical" operations will have "identical" results. Thus,
classifications are necessarily approximations or perspectives. But
we cannot do without them. This is something of a paradox: although
we know that strictly we are not allowed to consider things as
identical, we cannot generate knowledge if we don't. How can we
compare apples and oranges? By inventing the word "fruit". Only
mathematics says that we cannot compare them. In real life, we
compare them all the time, e.g. when we select an after dinner fruit
from the platter that is offered to us. The best we can do in
practice is to assume that things are "almost" identical and live
with the inaccuracies ("noise"), which may not matter that much.

     We say that one toss of a given coin is like another of the same
     coin, and that a result of heads is as probable as a result of
     tails. We say this because we have no reason to believe that any of
     the factors that matter to whether the coin lands on one side or
     the other have changed between tosses.

I don't think this is quite the situation. If all the factors affecting
the fall of the coin were the same, it would end up the same way every
time (almost). A supercomputer simulation given precise starting
conditions could probably predict the coin's final position with
considerable success.

The problem of a simulation is that it is only that: it incorporates
all of our (inexact) "knowledge" (obtained through "averaging"
previous results of the same class) about the laws that apply, but
this knowledge is just an approximation. And that in two ways: each
individual law will be inexact, and we won't know that we have taken
all applicable laws into consideration. In practice, only validation
of the supercomputer's program would indicate how good it is. We seem
to have a vicious circle: knowledge is generated through classifi-
cation; classification is not allowed; knowledge does not exist...

I think what probability means is an inability to predict an outcome,
because of a combination of imprecision in creating or measuring a known
starting condition, and a degree of complexity in the prediction itself
that is beyond normal human capacities to handle.

That is another part of it. An example of this can be found in
quantum mechanics, whose laws are very clear and pretty exact, yet
the computations are not easy. It is -- after a great many super-
computer hours -- possible to compute an electron's orbit around a
single proton with good (although far from perfect) accuracy, yet
accurate computations of orbits in larger atoms are still impossibly
time-consuming.

It is the lack of any relationship between the personal perception and
the actual determining factors that creates the notion of "chance."

I like that: chance is when I am not in control, whatever the reason.

The biggest objections I find to probabilistic treatments comes from the
use of a model that is obviously simpler than the observed processes it
is supposed to represent.

Then you will object to ALL applications of probability theory. Its
approach is to consider some things as "identical", because only then
can we average and predict some partial results with good accuracy,
although decidedly not everything to full accuracy.

Where you say this, you come remarkably close to my own position,
which is not to simply accept what others call "chance" or "dis-
turbance" but to try to model it as well as we can. Yet, somewhen we
must throw in the towel and work with inaccurate models, if only
because we have to control NOW, when our knowledge is still imperfect.

I think your argument begins to weaken when you try to establish a
relationship between subjectively estimated probabilities and
mathematically defined probabilities. As I pointed out above, the
essence of the concept of chance is our inability to predict or control
an outcome, but our refusal to blame that on our own ignorance.

There are, however, levels of ignorance. Although I cannot predict
the outcome of an individual coin toss, I can predict ABOUT coin
tosses -- as a class or category. One prediction is that I cannot
predict the outcome of an individual coin toss with better than 50%
accuracy. Another prediction is that, when repeated often enough,
"fair" tosses will cause "fair" coins to come down heads up in 50% of
the cases -- note the idealizations or classifications. Another
prediction is that I can predict no better. Thus, although knowledge
is, strictly considered, fully absent, meta-knowledge and meta-meta-
knowledge is there, and that pretty accurately. And these meta-levels
are pretty useful; they allow us to talk about averages, standard
deviations, habits, laws of nature, models, predictions, and a lot of
other useful things.

In mathematics, zero probability doesn't mean _a feeling of certainty_
that the event will not happen. It means that the event is treated as if
in fact it never happens. But subjective feelings of certainty most
often do not jibe with observations of frequencies of occurrance.

Sure. Humans are far from logical/rational. But studying statistics
can bring one a little closer. At least, it has helped me: I never
bet when I am not pretty certain that my odds are pretty good ;-).

Greetings,

Hans

[Hans Blom, 960705e]

(Bill Powers (960704.0700 MDT))

The "modern control theory" approach isn't, to my knowledge, concerned
with modeling actual human behavior, so the particular form of the model
that you have described hasn't been evaluated as a model of human
control processes, adaptive or otherwise. Am I wrong about that? If so,
I'd like to see the data.

A lot of pretty inconclusive experiments have been done in the
sixties and early seventies on how to model the "human operator", in
the hope that e.g. operators in a chemical plant's control room could
be replaced by control systems. But very little progress has been
achieved except, occasionally, by "expert systems", computer programs
that have formalized human knowledge built-in and that can "reason"
and come to conclusions and actions by "matching" the built-in
knowledge with the current measurements. But these systems are based
on heuristics and help little to advance our theoretical knowledge,
however useful they may be.

Where research is continuing, however, is in the analysis of multi-
input multi-output (MIMO) control systems. Not so much in
characterizing what goes on inside humans, but in coming to grips
with the complexity of these systems and in discovering "emergent"
laws. This is especially required when the system has more goals than
it has independent actuators. Then something akin to what Martin
calls "multiplexing" is required: dedicating the actuators to the
most urgent goals. You may notice this in a tracking experiment. When
the (Dutch) subject hears the word "coffee", he will very likely stop
tracking and fetch his mug.

So indirectly, the analysis of MIMO systems might throw some light on
what we call attention. But for the moment, little of specific value
to your interests is available.

I've spent most of today on catching up with my email, disregarding
the work I'm paid to do. Regrettably, my actuators do not allow me to
do everything at once. This has got to stop: it creates an immense
internal conflict ;- ).

Greetings,

Hans