[Martin Taylor 2015.04.27.09.52]
I don't know how it would be done in a neural system, and a computer
analogy can be misleading. But, without taking it too seriously as a
model, think of getting a value from memory in a computer. A vector
of ones and zeros is provided to the hardware, and a value pops out.
It’s such a standard operation that we just say that the value is
stored at such-and-such an address. I presume that hardware networks
are involved in extracting that value. But let’s look at it in
another way. There are lots of people in the world, and lots of
street addresses. We know only a tiny fraction of them, but they
exist.
We say that Mr. Jones lives at (vector) 234 Elgin Street, Mudville,
Yorkshire, UK. In this case, the vector is a successive refinement
of individually wide-ranging sets. In fact it’s a kind of matrix,
because you could say that “UK” refines “Elgin Street” of which
there are several scattered around the world, and there are lots of
addresses that are number “234” on their street. Where Mr Jones
lives is at the intersection of all these definitions. We can say
this because we know Mr Jones and have visited his house. We can use
“Mr. Jones” as an entry to a bunch of separate outputs for the
entire vector of the address, and we can use the address to provide
an output “Mr Jones” (and perhaps a few more of the family or other
residents).
Now think of it in a different way Mr. Jones may or may not be the
only person in the world who is about 187cm tall, weighs 152 kg, was
born April 1, 1983, has a scar on his middle finger right-hand, …
That also is a vector address for Mr Jones, and if someone had a
person with them and asked over the phone “Is this Mr Jones”, you
could say probably Yes or certainly No, based on this vector. But it
would be hard to use the vector to find Mr Jones in a crowd.
Come at this from a different direction, and think not of Mr Jones
and where he lives, but what would come out if one or more of these
intersecting activation elements changed value, say from “Elgin
Street” to “Lilac Avenue”, or from “April 1” to “June 22”. Would
anything come out? Would the new vectors bring up a different house
or person? Maybe, maybe not. If not, at least there would be some
activation from the other vector elements, and perhaps a bunch of
different memory units would have some low-level output.
So yes, networks must be involved, and entries that consisted of one
or a few elements of a vector would be expected to produce vectors
of several weak outputs, whereas entry vectors with many elements
would be expected to produce one or a few strong outputs.
It depends on the circumstances. If you want to see polaris in the
sky, then the associative memory is in the reference functions. I
was talking about it as being in the perceptual input, where you
have experienced “polaris” and “33W” together often enough (which
may be just once) that when you perceive one, the other is evoked. No working or simulated model of this, but we did use triflops (only
one of three outputs at a time is high) in hardware to run a series
of psychoacoustic experiments in the 1960s.
Do work though it, and notice that if the cross-link gains are low
you get enhanced outputs but no single one on each side of the
“labelling divide” is exclusively set high, whereas if the
cross-link gains are high, you get the flip-flop action on both
sides, and if they are very high you get rigidity and insensitivity
to contrary data. Also, note that if there are many elements on each
side (A, B, C, D, … and eh, bee, cee, dee, …) then a lot of
low-gain connections have the same effect (almost) as one high-gain
connection, which leads to what I said above: “entries that
consisted of one or a few elements of a vector would be expected to
produce vectors of several weak outputs, whereas entry vectors with
many elements would be expected to produce one or a few strong
outputs”, where “entries” in the diagram are the analogue inputs
from below and/or the values from the other side of the labelling
pair.
Again, only suggestions of possibilities. Could be on the wrong
track entirely, though to me, they just feel right, and should work.
A simulation might not be out of order to test that, but it’s not
one I will do very soon, as I am much involved in other things.
Martin
···
[From Rupert Young (2015.04.24 21.00)]
(Martin Taylor 2015.04.20.11.13]
Yes, that's what Bill was saying when he talked about
addressable associative memory. Depending on the address,
different memory values are called up.
Well, I am wondering how this is implemented in neural systems. If
the memory node (which is local to a control system) receives a
vector how is that converted into a specific single memory value?
What sort of structure would the memory node be? Sounds like it
may need to be network in itself (maybe like a hopfield-type
network, )
rather than an array on independent nodes which each encode the
individual memories.
So, in the example the address might be "polaris"
which gives a memory value of 33 W, say?
You have to ask what perceptions are controlled at the higher
level (that’s what the TCV is for) and see how they contribute
to the reference value for the perception you are interested in.
But in this example, it sounds more as though you are asking for
a perceptual input rather than a reference output. I know that
the distinction is rather murky when we are dealing with
imagination loops that produce imagined perceptions as a
consequence of output addressed access to memory, but the way I
look at it this morning (perhaps not yesterday or tomorrow) is
that you have heard someone say “polaris” and you have passively
perceived “33W” with control not yet involved. That’s what I
would call a “labelling” relationship.
I would have thought you would have a goal (reference) to perceive
“polaris”, and somehow this evokes the memory of “33 W” which is
then used as a gaze reference. So how do you get from polaris to
the 33W bearing? And how does the bearing control system memory
record all the other bearings, required for different stars?
Here's a very simplified sketch of the kind of circuit that
would do that. It’s not strictly hierarchic, and is therefore a
departure from Powers’s HPCT, but it would perform the function
and the results are controllable perceptions. This one
associates linguistic labels "eh’ and “bee” with visual forms
“A” and “B”. If “eh” is heard, it is likely that “A” will be
imagined, and vice-versa.:
A "flip-flop" is a circuit that tends to hold its outputs in
opposite “Yes-No” states because of positive feedback with
limits on the outputs of the “Amplifiers” as the continuously
variable input changes, until the input change gets large enough
to overcome the positive feedback, and the flip-flop flips to
the opposite state. You can make what I call “polyflop” circuits
in which there is one “Yes” state and several “No” in a group.
Long ago, we used such polyflops to control some psychoacoustic
experiments. The same kind of circuit could work on the output
side as well. If you want to perceive an “A” you might get it by
emitting an output “eh” to the person you are talking to.
Ok, looks interesting, do you have a working model? I'd need to
work through it to get to grips with it I think.
Suggestions, not assertions. And I don't know whether they are
in the direction of an answer to your original question.
Moving in the right direction, I think, though the waters be
murky!
Regards,
Rupert
On 2015/04/19 3:57 PM, Rupert Young
(
via csgnet Mailing List) wrote:
rupert@perceptualrobots.com
[From Rupert Young (2015.04.19 22.00)]
....
Not really, I am not sure what you are saying. Do you mean
that the local memory is a collection of different values,
and a specific value is retrieved by the vector address?
http://en.wikipedia.org/wiki/Hopfield_network