Model fitting -- Again.

[Martin Taylor 20020110 17:35]

Every few months I bring up the question of how best to fit models to
data. Each time, I think I'm getting a little closer to understanding
the problem--so here I am again.

Point 1: Something occurred to me that ought to have been
self-evident years ago: if the control system is built of linear
components (which includes integrators and differentiators), its
output can contain only frequencies that exist in the reference
signal or in the disturbance (apart from any natural oscillatory
frequencies inherent in its own structure).

It follows, then, that it is pointless to try to fit a linear model's
behaviour to that of a human outside the frequency band occupied by
the reference and/or disturbance (unless you are trying to fit a
natural oscillation such as is evident in Parkinson's disease). The
best-fit linear model must be the model that best fits the human data
_only_ within the bandwidth of the externally applied signals. This
is true no matter how complex and how many control levels are
involved.

If the human data contain frequencies outside the band of the
externally applied signals, it must be because of some nonlinearity
in the human. This includes such elements as data-dependent gain
changes, attention shifts (however they might be modelled, they must
involve changes in the gain of a variably-attended control unit), and
so forth. One should not attempt to fit those aspects of human
behaviour with a classical multi-level control system composed of
linear elements.

Point 2, which I have mentioned several times over the years: If the
human controls well, any control system that also controls well will
behave very like the human. The interesting aspects of the model fit
are the ways in which the human deviates from perfect control. One
possible "figure of merit" then is the accuracy with which the
model's deviation from perfect control matches the human's deviation
from perfect control.

As some of you are aware, I have, for some time, been trying to model
the behaviour of human pursuit trackers of arithmetically derived
targets and of graphically presented targets (i.e. they had to make a
user-controlled number equal to a continually changing result of a
simple arithmetic operation or make a user controlled line track the
length of an arithmetic combination of the lengths of varying lines).
The measure I have been using is the ratio of the RMS difference
between the user tracking error and the model tracking error to the
total user tracking error within the frequency band of the varying
target.

In doing this modelling, I have used not only the gain, lag, and
relative prediction gain parameters, but also a power-law
nonlinearity applied to the error signal before it is applied to the
output integrator. In the few examples I have so far tested, the
optimum power is usually between about 1.3 and 1.8, which suggests
that an inherent nonlinearity might exist. At least for the graphic
data, I probably should also do a logarithmic transform of the
perceptual input, but I haven't done that yet, and this might alter
the fit of the power law component.

I haven't heard of other model fits that introduced this kind of
nonlinearity into the connection between the comparator and the
output integrator, and I was wondering if anyone had tried it--and if
so, with what result?

Martin

[Kenny Kitzke (020111 8:15EST)]

<Martin Taylor 20020110 17:35>

<If the human data contain frequencies outside the band of the
externally applied signals, it must be because of some nonlinearity
in the human. This includes such elements as data-dependent gain
changes, attention shifts (however they might be modelled, they must
involve changes in the gain of a variably-attended control unit), and
so forth. One should not attempt to fit those aspects of human
behaviour with a classical multi-level control system composed of
linear elements.>

I suspect there is this non-linearity in humans, especially at the top of, or
around, the highest levels of perception by our body and mind sensors. Our
human spirit, including the matters of the heart of man, seem to not be a
linear (cause and effect) type circuit like we see in arm movements or
thinking.

IAE, if such non-linearities exist in humans, but not in the model, I take it
this would be one explanation for the limitation of the model to describe the
complete nature of humans and their behavior in the full PCT sense?

I wish I could get into your world to help improve the model as you imply,
but the nonlinear mathematics one would need are 35 years old in my cranium,
with not much application since.

Anyway, it is good to see someone working on the expansion of the science and
model we call HPCT.

[From Bill Powers (2002.01.11.0804 MST)]

(Ken Kitzke, 2002.01.11);

I suspect there is this non-linearity in humans, especially at the top of, or
around, the highest levels of perception by our body and mind sensors. Our
human spirit, including the matters of the heart of man, seem to not be a
linear (cause and effect) type circuit like we see in arm movements or
thinking.

Arm movements are not the result of a lineal, cause-effect, type circuit,
at least not in PCT.
"Linear," as Martin Taylor is using the term, means a causal relationship
described by a first-power equation, such as y = Ax + B. An example of a
non-linear equation would be y = Ax + Bx^2 + c, where the "^2" indicates a
superscript 2. The existence of a term with a power higher than 1 (2) makes
the relationship non-linear. Martin is proposing a power law relationship,
such as y = Ax^1.5. These causal relationships refer to properties of just
one element of the feedback loop, the output function.

To avoid this sort of confusion, I have recommended using the term "lineal"
to mean "in a straight line" or "sequential", and "linear" to refer to
curvature (or lack of it) in a causal relationship. A control system can
consist of _linear_ functions connected into a feedback loop, where a loop
is a _non-lineal_ organization whether its components are linear or
non-linear.

IAE, if such non-linearities exist in humans, but not in the model, I take it
this would be one explanation for the limitation of the model to describe the
complete nature of humans and their behavior in the full PCT sense?

If there are nonlinearities in real behavior but not in the model, I would
recommend putting similar nonolinearities into the model as Martin is
doing. This does not result in any basic change in the PCT or HPCT model.

Best,

Bill P.

[From Bill Powers (2001.01.11.0828 MST)]

Point 1: Something occurred to me that ought to have been
self-evident years ago: if the control system is built of linear
components (which includes integrators and differentiators), its
output can contain only frequencies that exist in the reference
signal or in the disturbance (apart from any natural oscillatory
frequencies inherent in its own structure).

True!

It follows, then, that it is pointless to try to fit a linear model's
behaviour to that of a human outside the frequency band occupied by
the reference and/or disturbance (unless you are trying to fit a
natural oscillation such as is evident in Parkinson's disease). The
best-fit linear model must be the model that best fits the human data
_only_ within the bandwidth of the externally applied signals. This
is true no matter how complex and how many control levels are
involved.

This is true also, but "bandwidth" is an elastic concept. In most real
systems, there is no sudden upper frequency cutoff; instead, at some
frequency the response begins to fall off at some rate -- a factor of two
per octave, for a system containing one integrator. This means that the
loop gain drops gradually as frequency rises, as well as entailing a phase
lag that increases with frequency. The "corner frequency" traditionally
used to measure bandwidth is the frequency where the amplitude response has
fallen to 0.707 of its low-frequency value. This means that the loop gain
is still half as great as its low-frequency value at twice the corner
frequency, and a quarter as great at four times the corner frequency.

In a simple integrating control system, the corner frequency of the output
function might be, say, 2.5 Hz, but with enough loop gain the closed-loop
response could be flat to 10 Hz or more. If there are significant lags, of
course, the loop gain is limited to smaller values and the closed-loop
response will be correspondingly limited (this is closely related to one of
your favorite subjects, the Nyquist sampling criterion -- I don't mean to
lecture you about something you know quite well).

If the human data contain frequencies outside the band of the
externally applied signals, it must be because of some nonlinearity
in the human. This includes such elements as data-dependent gain
changes, attention shifts (however they might be modelled, they must
involve changes in the gain of a variably-attended control unit), and
so forth. One should not attempt to fit those aspects of human
behaviour with a classical multi-level control system composed of
linear elements.

Agreed, although the amount of error remaining to be accounted for after
fitting to a linear model could be quite small. I'd say the importance of
introducing nonlinearities into the model would increase if the real
behavior showed striking departures from the model behavior.

Point 2, which I have mentioned several times over the years: If the
human controls well, any control system that also controls well will
behave very like the human. The interesting aspects of the model fit
are the ways in which the human deviates from perfect control. One
possible "figure of merit" then is the accuracy with which the
model's deviation from perfect control matches the human's deviation
from perfect control.

Yes, I've been saying that, too, for several decades. This is why I have
recommended using disturbances that are sufficiently difficult to produce
about a 10% RMS error in tracking studies. This makes control by the real
person deteriorate enough that a perfect control model with very high
integral gain and zero lag will not fit the behavior as well as a model
with finite gain and lag will fit it.

The way this shows up in fitting models to behavior is that for each
parameter, there is a best-fit value, with the fit becoming clearly worse
for values either higher or lower than the optimum. Perfect control
results in general from high gain and zero lag, so the model's fit becomes
worse as the parameters bring it closer to perfect control -- when the gain
is too high and the lag is too low.

If the difficulty of the control task is too low, the person's control will
come closer to perfect control, and the sensitivity of the fit to parameter
changes will become lower. With an easy enough task (small and slow
disturbance), the system noise swamps the changes of fit due to changing
parameters, and the best-fit model can't be distinguished from the perfect
control model. This may suggest the kind of "figure of merit" you're
looking for. The best-fit control parameters can be given error bars (say,
2-sigma), and the narrower the bars, the better the model.

In doing this modelling, I have used not only the gain, lag, and
relative prediction gain parameters, but also a power-law
nonlinearity applied to the error signal before it is applied to the
output integrator. In the few examples I have so far tested, the
optimum power is usually between about 1.3 and 1.8, which suggests
that an inherent nonlinearity might exist. At least for the graphic
data, I probably should also do a logarithmic transform of the
perceptual input, but I haven't done that yet, and this might alter
the fit of the power law component.

By "prediction gain" are you referring to first-derivative gain?

I've always thought that introducing nonlinearities would be a good idea,
and I still do -- congratulations on being the first (as far as I know) to
get out of his armchair and actually do it. The power-law approach is a
good one, giving you in effect a one-parameter nonlinearity to adjust. If
you still have residual non-random errors of fit, you might consider a
polynomial, in which you can adjust the contribution of different powers
individually. Second-power and third-power nonlinearities are especially
important -- symmetrical around zero, and non-symmetrical. Muscle
preparations, apparently, have an exponential force output response to
motor signal frequency, although a square law fits about as well.

I haven't heard of other model fits that introduced this kind of
nonlinearity into the connection between the comparator and the
output integrator, and I was wondering if anyone had tried it--and if
so, with what result?

As far as I know, you're the first. My Little Man model is very nonlinear,
but I haven't fit it to real behavior yet, and may never do so (a certain
amount of expensive instrumentation is required, not to mention an interest
in using it this way). I'll be most interested in seeing the results.

Incidentally, re model-fitting: In the advanced version of Vensim, there's
the ability to fit a model to data by varying parameters using something
called the Powell method. I tried it with some real tracking data which I
was already using to get model parameters by a method of successive
approximations. The Vensim method did _much_ better than mine; where mine
left a prediction error of around 5% RMS, the Vensim method brought that
down to less than 2% -- with the same data, the same parameter definitions,
and the same model! Clearly, it makes a lot of difference to use an optimum
method for fitting the model.

The nearest I've come to understanding the Powell method is to realize that
it starts with an n-dimensional grid which it searches for all the minima,
and then does some kind of curve-fitting (perhaps parabolic) to locate all
the minima as exactly as possible, and then picks the best one. It runs the
model quite a few times on the way to the result -- hundreds to thousands.
I found material on the Web under "optimization."

Best,

Bill P.

[From Rick Marken (2002.01.11.0900)]

Martin Taylor (20020110 17:35)

Great post Martin. Very thought provoking (which, by the way, is a
forbidden phrase if one believes that no one is responsible for what
anyone else does; "thought provoking" suggests that one is responsible
for -- one provokes -- the thoughts of another;-)).

Point 2, which I have mentioned several times over the years: If the
human controls well, any control system that also controls well will
behave very like the human.

I think it's important to point out that this is only true when the
control system that is used to imitate the human not only controls as
well as the human but also controls the same _perception_ that the human
controls.

In doing this modelling, I have used not only the gain, lag, and
relative prediction gain parameters, but also a power-law
nonlinearity applied to the error signal before it is applied to the
output integrator.

I think what this means is that you have made the error (e) to output
(o) relationship nonlinear (a power function), right? In other words, o
= e^p where the best fitting p is between 1.3 and 1.8. It's interesting
that the p parameter makes a big difference in the fit of the model to
the data. Since o = e^p is one form of the "universal error curve" your
results can be seen as a bit of evidence for that concept. (I hope this
doesn't cause apoplexy among the "antiuniversiterrians" -- the segment
of CSG that considers the universal error curve to be the height of
heresy;-))

Best regards

Rick

···

--
Richard S. Marken, Ph.D.
The RAND Corporation
PO Box 2138
1700 Main Street
Santa Monica, CA 90407-2138
Tel: 310-393-0411 x7971
Fax: 310-451-7018
E-mail: rmarken@rand.org