Anthropic developed a process to see what an AI is "thinking" about. They
found that when Claude is trained on text to predict tokens using a deep
neural network, that a short term memory emerges that is remarkably similar
to the way humans think and reason. Anthropic calls this global workspace
the Jacobian or J space, a set of neurons for each word that makes that
word more likely to be output.

For example, if I asked you the color of the fourth planet, you think
"Mars" and answer "red". When Claude does this, neurons corresponding to
"Mars" activate. Anthropic can read and modify these activations. For
example they can turn off "Mars" and activate "Earth" and Claude answers
"blue".

They can use this method as a lie detector for safety testing because words
like "fake" or "deception" will activate. The J space is used in less than
10% of processing. If the J space is turned off, then Claude can converse
mostly normally but it can't solve problems with intermediate steps and it
cannot lie because that requires thinking about what you really mean
without saying it.

The J space models the global workspace, the part of the brain that forms
conscious thoughts that communicates with all the other vital but
independent, unconscious processes like breathing, low level visual feature
detection, and controlling our 600 muscles in the right sequence when we
walk without thinking about it. It raises the question of whether conscious
thoughts require language, or are babies and animals conscious as well.

Anthropic did not design the J space. It emerged from training. A Jacobian
is the technique they use to detect it. Mathematically a Jacobian is a
matrix of partial derivatives of a vector of functions over vectors. The
paper introduction didn't go into details but I believe that Claude uses a
deep feed forward neural network (~100 layers) trained by back propagation
with multiple passes. Then the weights are frozen to prevent leaking
information between users. This requires a large context window because
after every question, Claude forgets the whole conversation and has to play
it back. The workspace emerges from the time delays going through a deep
network and a separate scratchpad memory holding a few dozen tokens.

But human brains don't work that way. Back propagation has no biologically
plausible mechanism. Instead, learning is local using Hebb's rule. Brains
have feedback loops, lateral inhibition, fatigue and time delays on the
order of 0.1 to 1 second. Our short term memory is only about 5 to 9
tokens, about half that of chimpanzees.

Hebb's rule only works one layer at a time, but natural language is
structured to be learned that way. We learn to recognize phonemes and
segment speech by 10 months before we learn our first word, then frozen by
age 6. In 2000 I found that you can restore deleted spaces in text using
only n-gram statistics with up to 77% accuracy for n = 5. This models the
tokenization process, which is hard coded in LLMs.

The introduction to the paper is here.
https://www.anthropic.com/research/global-workspace

Thank to James Bowery posting this link in the Hutter prize list.

-- Matt Mahoney, [email protected]

------------------------------------------
Artificial General Intelligence List: AGI
Permalink: 
https://agi.topicbox.com/groups/agi/T906c878e4ba411be-M019f5e6f40cd1aac567816b9
Delivery options: https://agi.topicbox.com/groups/agi/subscription

Reply via email to