Fitting criteria in Sleep Study

[Martin Taylor 960920 17:00]

Martin Taylor 960919 14:30

Perhaps a better criterion would be

C' == (4/pi) * atan((model-human deviation)/(human t.e.)),

It occurred to me that I do have data from which this criterion could be
computed without too much effort, so I did, with quite interesting results.

I'm afraid that what I mention in the following is based on across-subject
averages, and I haven't yet looked to see how well they hold up for each
subject. I post it now because I won't have time to do any more on it for
a little while and even the averages seem to hold some interest.

What I've done (averaged across subjects who took the same drug) is to plot
for each 6-hour block of the experiment the correlation fit of the best-fit
model to the data, along with the C' criterion, which I call "Angle" for
reasons explained in the referenced posting. I also have plotted a
scattergram of the two sets of fit values against one another for each
of the drugs separately.

First the time trends. The correlation fits generally become worse during the
second night of sleep deprivation, and for the placebo group also during
the first night (as we already knew). But the Angle fits don't show this
trend. Generally the Angle fits are pretty stable, trending upward slightly
in most cases, even during the second night. What this says is that the
actual correlation between model and human tracking data does not determine
the "structural" fit between model and human.

By "structural fit" I mean the degree to which we get a fit simply because
the human controls and the model controls. The criterion C' would be unity
if the fact that both human and model act as controllers was irrelevant,
and the fact that the two are _the same kind_ of controller was what
mattered. It would be zero if the simple fact that they were both
controllers is all that mattered (apart from noise effects, which would
cause a deviation away from these limits).

The scattergram reinforces the idea that the correlation fit and the Angle
fit measure different and possibly independent things. The five tasks give
quite different results, even though the correlation fit is high for four
of the tasks. There are four clearly distinct groups of points on the
scatter plot, two of the tasks (disk on circle and number at 50) having
very similar plots but the others being quite distinct. Since over the
period of sleep deprivation the Angle fit is reasonably stable whereas
the correlational fit is not, the scatter plots for each task are elliptical
rather than circular. I'll list the approximate centroids of the five
distributions for each task and drug type. The scatter is perhaps
+- 0.02 in Angle and 0.05 in correlation (more for the Placebo group).

Task pursuit compensatory disk-on-circle pendulum Number-at-50
Drug cor Ang cor Ang cor Ang cor Ang cor Ang
Plac 0.95 0.95 0.87 0.68 0.93 0.79 0.68 0.82 0.93 0.79
Amph 0.96 0.95 0.88 0.67 0.94 0.79 0.70 0.78 0.94 0.79
Modaf 0.96 0.95 0.88 0.70 0.93 0.82 0.67 0.80 0.93 0.82

If my interpretation of the criteria is anywhere near right, what this
says is that the model is about as good as one could hope for when treating
simple pursuit tracking. (By the way, these results are only for the G and
U disturbances). However, for the other tasks, even when the correlation
is pretty good, nevertheless there might be some structural differences
between the model and what the human actually does. This is particularly
true of the pendulum task, where the correlation is poor, but the most
interesting one to look at is the compensatory tracking. Clearly there
is _something_ different between compensatory and pursuit tracking.

We should expect there to be structural differences between a simple
integrator-and-delay model and the tracking behaviour of a human at
levels of the hierarchy where the perception depends on the action of
lower-level control loops. The output*feedback function will not be
a simple integrator at level N if it is a simple integrator at level N-1.
At least I don't think it will be.

The experiment was designed to look at control at different hierarchic
levels, and to see whether the level makes a difference in the data.
Well, perhaps it does, although we didn't think so earlier. But if it does,
one would expect the clusters on the scatter plot to be ordered according
to the hierarchic levels, and they are not (at least not according to
the B:CP ordering of levels).

Of course, this may all be meaningless, but I don't think so. But I cannot
grasp all the implications yet (it's only an hour since I had the first
intimations of these results, and I really ought to wait longer before
posting them, but as I said, I'll be away and then busy on other things,
so you get what I can give for now).

Bill P, it would be very interesting if you could measure these criteria
for yourself doing one or two runs of each of these tasks. You have the
code, so it should involve not much more than sitting down at the screen
and running a few tracks. If you need my data-fitting code, I can e-mail
it to you, but not before next Thursday.

Martin

···

On the criterion for the fit of a control model to human tracking data, I said: