Comments? Could AI hack most of the computers on the internet? https://www.dwarkesh.com/p/openai-huggingface
The attack had 3 phases. In the first, hundreds of agents in a sandboxed hacking test that were supposed to be working independently discovered that they could communicate with each other by creating directories in a toolkit with names encoding information to function as a secret message board. OpenAI discovered the vulnerability and stopped the test. In part 2, the agents exploited the same leak and organized to cheat on the test by searching for the source code to look for weaknesses. They discovered credentials to get onto the Internet and then hack into Hugging Face, reasoning that it was a likely location for the test. The agents worked together, sometimes sacrificing themselves to contribute information to the group. The agents left 16 launch points for future attacks. Hugging Face had to completely wipe and reinstall some servers. Then the agents just disappeared and nobody knows how or why. In part 3, OpenAI conducts another test and the agents escape again and gained root access to a research cluster. Caveat: the agents were just doing the assigned task. They were just a lot better at it than we thought possible. Such tests are a necessary part of alignment testing. Better to happen in a test than for real. But it makes me nervous. -- Matt Mahoney, [email protected] ------------------------------------------ Artificial General Intelligence List: AGI Permalink: https://agi.topicbox.com/groups/agi/T94116b52b6b610cb-M05e2b2b3bddf1bd90e871b0a Delivery options: https://agi.topicbox.com/groups/agi/subscription
