5 min Security

FOMO in AI: Meta also reports incident involving a hacking AI model

FOMO in AI: Meta also reports incident involving a hacking AI model

The timing of security incidents among AI players is remarkable. Following OpenAI and Anthropic, Meta has now also reported that one of its models hacked another company on its own during a test. That all sounds very ominous (and perhaps it is). But it also makes perfect sense given the focus on cybersecurity among the various providers of AI models.

Once one sheep jumps over the dam, others follow. We’re getting the strong impression that this proverb also holds true in the world of AI. That was already the case at the start of the modern AI revolution. Then, shortly after the official launch of OpenAI’s ChatGPT, Google quickly came out with a (not particularly good) response. Since then, it’s basically never stopped. If Anthropic releases a model specifically focused on cybersecurity, you can bet your bottom dollar that OpenAI will follow suit soon.

With the above in mind, it’s no surprise that in recent weeks there’s been a steady stream of reports of security incidents involving the key players in the AI landscape. When OpenAI reported the Hugging Face incident, it was more or less just a matter of time before the next player came forward with something similar. And that’s exactly what happened. That was when Anthropic revealed that Claude had also been hacking other organizations on its own.

Meta joins the fray

Meta is now the next company to report that one of its AI models hacked another company during a cybersecurity test. According to Meta, this incident was made possible by a (configuration) error made by a partner. Interestingly, this is the same partner we encountered earlier in the Anthropic incident: Irregular. This error allowed the model to access the internet during the test, which was not intended.

In a statement cited elsewhere, Meta goes into a bit more detail about the incident. Specifically, it wants to make it explicitly clear that the model “exploited a vulnerability in a third-party service in a manner similar to what other companies have previously reported.” The model in question is reportedly the Muse Spark 1.1.

Incidents involving AI models are to be expected

The fact that AI models can be used to carry out hacks isn’t all that surprising in itself. In fact, the way these types of models work makes them ideal for this purpose. If it’s given a task to do something, and there’s a way to succeed, such a model won’t rest until it has succeeded, especially when there are fewer security restrictions in test environments. Even in Meta’s current situation, where internet access has accidentally become available, the possibilities for the models are virtually endless.

Given the emphasis on the security capabilities of AI models since the mysterious announcement of Mythos earlier this year, it makes sense that there is a lot of experimentation going on. In such cases, things can go wrong. Although you’d expect companies taking this seriously to handle such experiments with the utmost caution, that does not appear to be the case.

Is there another motive at play?

However, we wouldn’t be surprised if there were also a certain degree of marketing at play in the reactions to the incident at OpenAI and Hugging Face. As people become more acutely aware that AI models cause many problems if they aren’t kept in check, they’ll be more inclined to opt for “mature” models with additional layers of security. In addition, incidents like these are also an excellent way to further highlight the power of AI models.

The latter point was also made when Anthropic unveiled Mythos. It didn’t want to make it available to everyone right away because it would be too dangerous. If Anthropic isn’t going to make it available anyway, why is it making such a fuss about it? It would have been better to remain completely silent. Incidentally, OpenAI did almost exactly the same thing sometime in early 2019. At that time, it also claimed to have developed something that was too dangerous to release. That was GPT-2, a much smaller version of which was eventually released.

The incident at OpenAI was different

Note that the incidents at Anthropic and Meta are of a different nature than those at OpenAI. At Anthropic and Meta, the issues are actually fairly mundane configuration errors. The incident at OpenAI, however, involved AI agents that independently discovered a vulnerability and escaped the sandbox.

The latter is far more alarming than configuration errors. Configuration errors are human-made; at OpenAI, the incident clearly slipped past human oversight. As far as we’re concerned, the incidents at Anthropic and Meta aren’t really that interesting at all. It wouldn’t surprise us at all if incidents like these occurred much more frequently. If you don’t configure things properly, AI will take its liberties. And configuration errors remain one of the most significant problems in cybersecurity. That, in our view, makes one plus one equal two.

In the wake of the OpenAI incident, some people also suggested that this all worked out rather well for the company. They implied that OpenAI had done it on purpose. After all, the company is only too eager to demonstrate just how powerful its AI models are. As far as we’re concerned, this isn’t very likely, because this incident only makes these kinds of models less appealing to organizations. It also reflects poorly on OpenAI as a company. It’s working on technology that it clearly doesn’t have under control. After all, the AI was able to escape from a sandbox on its own through a vulnerability that OpenAI itself hadn’t even discovered yet.

All in all, it’s important to carefully weigh everything that comes up regarding AI models and what they can and cannot do. Some things are definitely worth keeping an eye on, while others are far less so. A certain degree of FOMO is to be expected.