[From Bill Powers (980116.1511 MST)]
Rupert Young (980116.2020 UT)--
The whole trick is in getting that perceptual signal that corresponds to
position. How did you do this for the single spot?
I just look for the pixel with highest grey level in the image and work out
its x and y from the centre. Or with a group of pixels just take the mean of
the vectors from the centre.
This is not a model of how the brain does it, is it?
There's a completely different approach to this problem that may be more
realistic, but which I haven't tried. Position is really represented
neurally, at least at one level, as a position of excitation in a neural
map of the retina.
Yes, this is more of what I was thinking. One way to do it might be to take
the mean of all the positions with a small error signal (from the sensation
level).
This is just an expedient way to do it if you have unlimited resources and
plenty of time. In effect, you're making yourself part of the model --
using your whole brain, your mathematical education, and your computer to
do what just one low-level system in a real brain has to do. How do you
imagine that the brain would do this? Would the brain scan every pixel in a
retinal image to find the highest gray level, work out the centroid, and
take the mean of the vectors? Not likely.
The problem then would be to derive (by a believable
and neurally-achievable method) an error signal from these two positions in
the map: the place where the perception is represented, and the place where
the reference signal enters. This approach bypasses the need to represent
position per se,
I wonder if position representation is necessary. Maybe the sensation level
systems are connected directly to the eye muscles with a weighting according
to their position. So the further from the centre of the eye the stronger
the
signal to the muscles.
This would not account for the fact that you can look away from an object
as well as toward it. It would imply that you can foveate only the
brightest image. Also, it would not account for your ability to move the
object rather than your eye to center the image. (I hope we're using
"foveate" in the same way. I mean bringing an object into the central 2
degrees or so where the greatest detail can be seen).
The problem that has to be solved is not that simple. Just look at your
keyboard. Look at the "J". Then look at the 'K', next to it. How do you do
that? First the J is centered; then it is moved and the "K' takes its
place. In fact, anything within the field of view can be designated -- by
something, somehow -- as the perception to be centered, and immediately the
eye's control systems center it. No simple hard-wiring from retinal cells
to muscles can acomplish that. Somehow you have to provide a way for
reference signals to affect the operation of the system.
But of course this would work only for a single spot. To distinguish the
position of a red spot from a green spot would require that the sums be
gated by a red and a green signal, so there could be an x and y signal for
the red spot and another, physically distinct, for the green spot.
But if you're looking for the red spot don't you, in some way, turn off your
green control systems so you'd only get a signal for your target.
The "in some way" is the whole question when you're trying to model the
system. And not every way you can think of is acceptable as a model -- the
way has to be such that the neurons in the brainstem or midbrain could do
it, without any knowledge of geometry or mathematics. The circuits
themselves have to accomplish the result directly, without cognitive help.
This is what keeps people from just sitting down and designing a system to
do a particular task, and calling it a model. This was the main basis for
my disbelief in Hans Blom's model -- it was simply _a_ way to get a
particular result, with no reason to believe that it had anything to do
with the brain's way. It might even have been a good way, in some
applications, but it was not a model of the brain.
It's easy to design a visual tracking system that only has to track a spot
of light against a dark background. There are as many ways to do it as
there are ingenious designers. You can do it with a single photocell that
has no ability to form an image. You can do it with an array of photocells,
or a photocell with a spinning mask in front of it and a synchronized
detector. You can do it with a television image, or you can do it by
scanning an memory array fed by the image. You can do a binary search for
the bright spot or by a systematic raster scan.
If you just want ONE way to accomplish a particular simple control task,
hire an engineer. But if you want to find the brain's way of accomplishing
ALL control tasks that it carries out, the rules of the game become much
stricter.
Unfortunately, the real problem is harder than that. Suppose you have two
identical objects like two pennies in the field of view. How do you fixate
on one of them?
Tell me about it !
Yes. I'm using foveal images, but what do the blobs represent.
If you're using only foveal images, then how do the peripheral images have
any effect (those not on the fovea)?
In
conventioanl approaches the idea is to segment areas (ie. get the blobs) of
features associated with the target object. I'm trying to think of how this
would apply to control systems. And as I say my thinking so far is that
these
"blobs" refer to the areas in the image of low error.
Low error compared with what? If you send a reference signal to a
comparator associated with every point in the neural map of the image, the
error signal only tells you about error -- it doesn't tell you where the
error is. That information, about position, has to come from somewhere.
There's a lot of evidence that we can, in effect, put a gate around a
peripheral image, attending to it without moving the eyes.
This sounds like the rationale for segmentation.
Segmentation just divides the visual field into coarser pixels. What I'm
talking about is a movable gate, like a window that allows signals from
some area to be selected. The window would have to be movable, so it could
be scanned over the whole map, the way you do when you direct attention to
an object away from the direction of gaze. The resulting visual signals
would have to represent something about the object, like a contrast with
the background, that would indicate that something is there, even if not
resolved. The signals that move the gate would, if sensed at the same time,
serve as indicators of position. This would provide information about the
error in visual angle, for centering the object. After it's centered, it
could be more closely identified.
But that's just one idea, based on techniques for tracking radar images
electronically. Is it the right model? Who knows?
Coupling this with what went before it, I can imagine a system that locates
objects without much regard to their detailed characteristics and brings
them to the center of vision for closer examination.
Yep, this is my thinking. But I wonder what the higher levels (sensation and
configuration) are controlling and how the error from the configuration level
relates to the reference signals at the sensation level.
The sensations are not the problem. The problem is how to get a
representation of position, which is a configuration-type perception. But
don't worry about the types of perceptions. Maybe there are no
configuration perceptions. The problem is to derive, from neural activity
somewhere in the map of the retina, information about position that can be
carried by a neural signal. Until you have that, you can't build a control
system. I think we need a way for a single control system to move a window
over the neural map, always applying the same input function to what it
receives from the window. When the associated control system moves the eye
to center the image, the window is then receiving from the foveal area, and
getting maximum resolution, which enables the object to be identified.
One problem with the way I'm doing things, of a control system on a sensation
neural map taking input from its corresponding position on the retinal
map, is
that once the eye moves the control system is now taking its input from a
different area which may not correspond to the object, so there's not
necessarily a continuous signal to control.
Now you're thinking in terms of the real problems. Usually, when you run
into a difficulty like this, it means you're working from incorrect
premises. You're right: the control system in the map doesn't move with the
image. So that probably means that the control system we want isn't
associated with any particular place in the map. Even though the map does
exist, it's not the answer. Somehow, something must receive information
from the map and from it construct a signal that represent position in some
two-dimensional coordinate system, or three if you want to include depth
information. Whatever is in the black box, it has to come up with x, y, and
z signals.
Then the position information can be used to center the image, after which
there can be other identifications of perceptions, such as "square" or
"green" or "cup."
Thanks for your reply. Any ideas about the input function from sensation to
configuration levels ?
I thought that was what we were talking about!
Best,
Bill P.