Comments? Could AI hack most of the computers on the internet?

https://www.dwarkesh.com/p/openai-huggingface

The attack had 3 phases. In the first, hundreds of agents in a sandboxed
hacking test that were supposed to be working independently discovered that
they could communicate with each other by creating directories in a toolkit
with names encoding information to function as a secret message board.
OpenAI discovered the vulnerability and stopped the test.

In part 2, the agents exploited the same leak and organized to cheat on the
test by searching for the source code to look for weaknesses. They
discovered credentials to get onto the Internet and then hack into Hugging
Face, reasoning that it was a likely location for the test. The agents
worked together, sometimes sacrificing themselves to contribute information
to the group. The agents left 16 launch points for future attacks. Hugging
Face had to completely wipe and reinstall some servers. Then the agents
just disappeared and nobody knows how or why.

In part 3, OpenAI conducts another test and the agents escape again and
gained root access to a research cluster.

Caveat: the agents were just doing the assigned task. They were just a lot
better at it than we thought possible. Such tests are a necessary part of
alignment testing. Better to happen in a test than for real.

But it makes me nervous.

-- Matt Mahoney, [email protected]

------------------------------------------
Artificial General Intelligence List: AGI
Permalink: 
https://agi.topicbox.com/groups/agi/T94116b52b6b610cb-M05e2b2b3bddf1bd90e871b0a
Delivery options: https://agi.topicbox.com/groups/agi/subscription

Reply via email to