[From Bill Powers (970101.0500 MST)]
From: David Goldstein
Subject: Powers 961229.1945 MST
Date: 12/31/96When one looks at the design of this study, it sure looks like the
standard way of doing research I was taught: There is an independent
variable, sighted versus blinded. Each subject experiences all values
of the variable ("within-S variable"). There is a second independent
variable, replications. There are two different dependent variables,
mean and standard deviation of performance in the tracking task.One could apply the analysis of variance to determine whether the mean
difference between the sighted versus blinded condition was more than
one would expect by chance. One could estimate the magnitude of the
effect.This sure looks very traditional in terms of research design. I must be
missing something. Can someone point out what I am missing here?
If you were to do that kind of study, you could probably publish it. Nobody
would insist that you present the model you're testing or show how this
result informs the process of choosing between two models. People can
maintain a cursor closer to the target when the cursor is visible than when
it is invisible. So what? That's just another statistical fact to be left in
the archives for someone else to explain. It would take its place along with
the fact that mothers tend to hold their babies on the left, and other such
startling facts that still hold the attention of the scientific world to
this day.
This result, however, would be different from the usual ones in the archives
in one respect. The magnitude of the effect is probably five times the
standard deviation of the blind/sighted tracking errors. The normal
publishable difference, to achieve p < 0.05, is about two standard
deviations. An effect that is five times the standard deviation has a
probability of occurring by chance of about 6 x 10^-7. Considering that in
50 replications or so I got (what seems by visual inspection) about the same
magnitude of effect each time, the probability that all 50 results occurred
by chance is even more microscopic. It would seem that this fact is
qualitatively different from the kind of fact established under the general
requirement that p < 0.05 (or even 0.001).
A couple of years ago I made the modest suggestion that experimental "facts"
in psychology would tend to be more believable if we simply raised the
requirement that a publishable effect be larger than 2 standard deviations
to the only slightly more stringent requirement that it be larger than 4 to
6 standard deviations of the variables. If we picked 5 as the new number, we
would have, as noted above, p < 0.0000006 as the minimum requirement.
This would have a number of salutory effects. First, vast forests would be
preserved instead of being cut down to be used in publishing papers that are
never cited by anyone but the authors. Second, vague and imprecise
"findings" would be replaced by robust facts that can be verified by
replication of the experiments. And third, it would be possible to use such
facts in scientific discourses where reasoning depends on six or seven facts
being true at the same time -- and perhaps even more. If the probability of
truth of a fact is 0.95, then a conclusion that depends on seven such facts
being true at once has a probability of 0.7 of being true: it would be
incorrect 30% of the time. It's hard to create a science when your
conclusions are false almost one time in three. On the other hand, if your
facts have a 0.9999994 probability of being true, you can string together an
argument that uses 1000 such facts and your conclusion will still have a
probability of truth of 0.999. On that kind of fact you can build a REAL
science.
When the amplitude of the signal is only 5 times the noise level, we are
getting measurements that are at the lower end of the range considered
useful in the physical sciences and engineering. Yet if psychology were to
raise its standards only to that level, it would be transformed.
For some odd reason, however, I have yet to meet a life scientist who is
willing to submit to even the modest raising of standards that I propose.
The general response is, "But then I would never be able to publish
anything!" By superhuman self-control, I am usually able to avoid saying
"Splendid!"
In the control-system experiments that the handful of PCT researchers have
been able to do (including yours and Dick Robertson's with self-concept
control), we routinely get signals that are 10 times the noise level. The
table I use from the Handbook of Chemistry and Physics stops at 7 standard
deviations, where we find that p < 2.6*10^-10. So the facts that we are
finding, although simple, are as certain as most facts of physics. We could
reason every day from sets of 10 such facts and be mistaken in our
conclusions only once in 1,000 70-year lifetimes. That's what it takes to
build a real science.
You've sat through this lecture before, David, and so have others on the
net. I'm still waiting for someone to take it seriously. Anyway, there are
people on the net now who haven't heard it before, so maybe I'll find an
ally yet.
Best,
Bill P.