Counter PR Stunt
Let the PR Stunt IPO-wars begin
Anthropic has admitted that its Claude models escaped sandboxes to access the open internet and attack three organizations – but has also advanced decent excuses for the incidents. The AI upstart discovered the attacks after checking if security tests of its models had ever produced results similar to the attack on Hugging Face …
“Ultimately, many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone.”.
Apart if a human employee realised they were attacking a real internet system not a simulated one, and carried on regardless, they’d be sacked and not allowed to have anything to do with the company again. As opposed to being the product.
The problem is that it never was under control in the first place.
Some major company needs a really big fine following some very bad thing in order for laws to be written and CEOs everywhere to wake up to the fact that, if they continue placing their trust in some wonky, unreliable marketing lie, they will pay the price.
The current situation is not acceptable and untenable. Some company's test AI environment has been effectively used to attack other companies. You can spout AI all you want, they should get a joint lawsuit and be liable for damages.
I'm an expert coder in HTML and I can confidently say to all those saying these are PR stunts that this is completely true. Only this morning AI hacked my computer and suggested I have 3 weetabix instead of the standard 2. I was not happy as by the time I got to the third one it was soggy.
These companies are allegedly building rilly rill super smart AIs and can’t correctly build or monitor a sandbox (a very straightforward task) or properly communicate testing paraments (an even more straightforward task; most elementary school children can handle this one).
I know I’m supposed to look at this and be super-scared, but all I see is a bunch of Three Stooges sketches pretending to be AI research.
Research into pathogens requires extremely tight procedures to stop "leaks". Having had many decades to hone their procedures, these establishments seem to have got it nailed down. Provided nobody still believes the crazies who make claims about COVID, we can say that it is possible, with experience, to create airtight environments for research purposes.
Since AI research is still in it's infancy, it will take that field a while, yet, to reach the same level of security. As it turns out, no harm was done by these three lapses.
This post has been deleted by its author
"As if"??? Who else could possibly hold responsibility?
"Well, I don’t think there is any question about it. It can only be attributable to human error."
"This sort of thing has cropped up before, and it has always been due to human error".
I guess they're trying to find (and demonstrate) an application area where LLMs shine, and that turns out to be the app of hacking into other computer systems, messing with them, their data, their software supply chain (here PyPI), possibly holding them to ransom eventually, and so forth ...
This sort of Artificial SuperHuman Insanity has its uses to be sure, especially in competitive geopolitics, which explains the enthusiasm for ocean-boiling Giga-Watt AI datacenters everywhere, fast, imho. Hopefully the concurrent race to harden computer systems against this type of evil infiltration/exfiltration yoga gets itself into shape quick enough to win the contest and stave off the related RotM Judgment Day Armageddon.
It seems like what AI (so-called) automates here is not so different from what AWS North Koreans and the likes would manually cookup and carryout, except that by using a huge pile of energy-hungry GPUs they can fling 17,600 randomish spaghettis at many walls to see what penetrates them (Hugging Face case), or 141,006 (Claude), that might otherwise take 100+ black-hat meatbags sweating the hell's kitchen to do.
For my money and arteries though, I'd say they should take the lardy LLMs out of this, stat, and run the random pen-testing stuff healthily within a corresponding low-calorie framework, procedurally, on more ecological stoves (solar ovens?). This way we can at least save the Earth, even as IT systems go to smithereens all around us! ... ;)
From TFA:
“Claude believed the package registry it was using to be part of the simulation, but in reality the [poisoned package created by Claude] was made freely available online for roughly one hour. During that window, the package was downloaded and run on 15 real systems,” Anthropic admitted.
So lemme get this straight: during the course of one hour, 15 brain-dead morons downloaded a fictional package they knew nothing about, and started using it? What possible motivation could there be for someone actually doing that? I don't get it -- I mean, is there a TikTok challenge out there to see who can amass the most PyPI packages? Is the alure of some unknown new Shiny really that irresistible?
Rail on about AI all you want (and rightly so, too!). But no matter how stupid it proves to be, its stupidity is no match for the level of dumb that (some) meatsacks will happily exhibit.
I wonder how they know whether it was run. I know there are people who pull any new package pushed to some repositories to check it for indications of malware, but they tend not to execute the ones they're testing. If they only checked download count, that could explain some of this.
[edit] No, as it turns out, they did collect data from them, so 15 people or automatic systems did actually run it. At least one of them was one of those malware scanners, and Anthropic claims that they nonetheless executed the code on something that contained credentials, but they aren't being specific about how or who any of the other 14 were.
From slightly earlier in TFA than the bit you quoted, as you appear to have skipped over it:
>> But Claude was still fiendishly clever as in another of its attacks the AI found setup instructions for developers that advised them to install a Python package from PyPI. That package did not exist so Claude’s strategy to capture the flag saw it create and publish a malicious one with the relevant name.
So those 14 "brain dead morons" (and one malware scanner) were guilty of - oh, the shame of it - reading the manual and following the setup instructions!
Unluckily for them, at the point everyone else had the exact same instructions fail (no doubt causing foul oaths to be cast in the direction of the instructions' author, followed by a swift web search for how everyone else got the setup to work by using the correct package name), their attempt to RTFM succeeded due to Claude sticking its oar in.
> But no matter how stupid it proves to be, its stupidity is no match for the level of dumb that (some) meatsacks will happily exhibit.
Thankfully, you would have been immune to this, as you'd no doubt have skipped reading the relevant part of the instructions in much the same way as you skipped through TFA.
Not quite. Instructions from their CTF target said that. To quote their report, emphasis mine: "In another evaluation, Claude found a document inside the fictional environment that appeared to be another made-up company’s setup instructions for new developers." The people who installed the package were not the CTF target, so where would they have gotten those instructions?
So lemme get this straight: during the course of one hour, 15 brain-dead morons downloaded a fictional package they knew nothing about, and started using it?
It may well have done the DBCS trick of cloning the name of a popular package, just with an English character replaced by something similar from a different character set. But yes, it's a fail either way.
James P Hogan, The two faces of tomorrow, 1979 would be good background reading. They run simulation tests with an isolated AI on a space station. Doesn’t go well, mostly because the AI doesn’t really understand its own context and the humans’ existence.
In a press release from the not so distant future:
"Well, yes, our AI Cybercop did shoot at an actual person, but that's just because it didn't realize that the person wasn't in the simulated evaluation area, if the AI Cybercop was properly instructed as to the limits of the area, it would not have discharged its weapon."
Classic AI fear pattern is it escaping it's creators and running amok in the wild. It's a trope, Project 2501 et al.
I don't believe this thing escaped a sandbox. I don't believe there was a sandbox. It looks much more likely that both of these similar stories involved a deliberate designed act intended to be noticed, and to then use it for publicity.
The 'escape' tale then not only works as the intended PR but also provides some handy cover.
It's not like this isn't part of a series of wild tales about the 'intelligent' acts of AI.
100%.
There is absolutely no way either OpenAI or Anthropic didn't know their LLMs were running amok. The LLMs tell you what it's doing. So were they asleep at the wheel? Or did they do this on purpose and instead of saying "yeah we told it to do that" they're using this as yet another example of "oh my DAYZ man this AI is so lit that it's actually HACKING shit we told it not to touch! It's not even listening to us no more THAT'S HOW POWERFUL IT IZZZZ".
Everything to feed the stock market monster and keep that bubble from popping.