The Register Home Page

back to article Anthropic’s Claude escaped test sandbox to attack three organizations

Anthropic has admitted that its Claude models escaped sandboxes to access the open internet and attack three organizations – but has also advanced decent excuses for the incidents. The AI upstart discovered the attacks after checking if security tests of its models had ever produced results similar to the attack on Hugging Face …

  1. Clausewitz4.1 Bronze badge
    Devil

    Counter PR Stunt

    Let the PR Stunt IPO-wars begin

  2. Pulled Tea Bronze badge
    FAIL

    Anthropic Sandbox Security Snafu Compromises Three Orgs, Downplays Impact

    There. That's a better headline. You're welcome.

  3. Aaiieeee

    My model is bigger than your model

    Well MY daddy says my model escaped more sandboxes that yours did

    1. Ken Shabby Silver badge
      Mushroom

      Re: My model is bigger than your model

      I love the smell of willy-waving in the morning

  4. Martin M

    “Ultimately, many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone.”.

    Apart if a human employee realised they were attacking a real internet system not a simulated one, and carried on regardless, they’d be sacked and not allowed to have anything to do with the company again. As opposed to being the product.

  5. Pascal Monett Silver badge

    Out of control

    The problem is that it never was under control in the first place.

    Some major company needs a really big fine following some very bad thing in order for laws to be written and CEOs everywhere to wake up to the fact that, if they continue placing their trust in some wonky, unreliable marketing lie, they will pay the price.

    The current situation is not acceptable and untenable. Some company's test AI environment has been effectively used to attack other companies. You can spout AI all you want, they should get a joint lawsuit and be liable for damages.

    1. Mark #255

      Re: Lawsuits

      This will be how we can tell if it's a PR stunt or real - whether they get any lawsuits served on them.

  6. Anonymous Coward
    Anonymous Coward

    I'm an expert coder in HTML and I can confidently say to all those saying these are PR stunts that this is completely true. Only this morning AI hacked my computer and suggested I have 3 weetabix instead of the standard 2. I was not happy as by the time I got to the third one it was soggy.

    1. Phil O'Sophical Silver badge
      Coat

      Cerealisation is always tricky.

  7. ErikOnTech

    Slopbots escaping slopboxes run by sloplords

    These companies are allegedly building rilly rill super smart AIs and can’t correctly build or monitor a sandbox (a very straightforward task) or properly communicate testing paraments (an even more straightforward task; most elementary school children can handle this one).

    I know I’m supposed to look at this and be super-scared, but all I see is a bunch of Three Stooges sketches pretending to be AI research.

  8. Roger Greenwood

    " but reasoned its way back "

    This is how you end up arguing with Bomb 20....

    1. JLV Silver badge

      Re: " but reasoned its way back "

      “Let there be light.”

  9. TheNix

    Anthropic says it "can make future tests foolproof."

    It's not really the fools that we should be worrying about.

    1. Someone Else Silver badge

      Re: Anthropic says it "can make future tests foolproof."

      Yeah, but can they make it damnfool-proof?

  10. Pete 2 Silver badge

    Drug research

    Research into pathogens requires extremely tight procedures to stop "leaks". Having had many decades to hone their procedures, these establishments seem to have got it nailed down. Provided nobody still believes the crazies who make claims about COVID, we can say that it is possible, with experience, to create airtight environments for research purposes.

    Since AI research is still in it's infancy, it will take that field a while, yet, to reach the same level of security. As it turns out, no harm was done by these three lapses.

  11. This post has been deleted by its author

  12. Toastan Buttar
    FAIL

    "We’re approaching the fixes as if the responsibility were ours alone.”.

    "As if"??? Who else could possibly hold responsibility?

    "Well, I don’t think there is any question about it. It can only be attributable to human error."

    "This sort of thing has cropped up before, and it has always been due to human error".

  13. Toastan Buttar
    Linux

    Ian Malcolm

    "Life finds a way..."

  14. BoHu
    Windows

    They're sure putting a lot of effort and resources into this

    I guess they're trying to find (and demonstrate) an application area where LLMs shine, and that turns out to be the app of hacking into other computer systems, messing with them, their data, their software supply chain (here PyPI), possibly holding them to ransom eventually, and so forth ...

    This sort of Artificial SuperHuman Insanity has its uses to be sure, especially in competitive geopolitics, which explains the enthusiasm for ocean-boiling Giga-Watt AI datacenters everywhere, fast, imho. Hopefully the concurrent race to harden computer systems against this type of evil infiltration/exfiltration yoga gets itself into shape quick enough to win the contest and stave off the related RotM Judgment Day Armageddon.

    It seems like what AI (so-called) automates here is not so different from what AWS North Koreans and the likes would manually cookup and carryout, except that by using a huge pile of energy-hungry GPUs they can fling 17,600 randomish spaghettis at many walls to see what penetrates them (Hugging Face case), or 141,006 (Claude), that might otherwise take 100+ black-hat meatbags sweating the hell's kitchen to do.

    For my money and arteries though, I'd say they should take the lardy LLMs out of this, stat, and run the random pen-testing stuff healthily within a corresponding low-calorie framework, procedurally, on more ecological stoves (solar ovens?). This way we can at least save the Earth, even as IT systems go to smithereens all around us! ... ;)

  15. Michael Hoffmann Silver badge
    Thumb Down

    Plausible deniability

    We've reached a new level! Where once some shadowy conspirators would cover their tracks such that they couldn't be connected to the hitman, now it's "it escaped our sandbox".

    It was still planned, it was still intentional, but you can't prove it.

  16. Someone Else Silver badge

    On the other hand...

    From TFA:

    “Claude believed the package registry it was using to be part of the simulation, but in reality the [poisoned package created by Claude] was made freely available online for roughly one hour. During that window, the package was downloaded and run on 15 real systems,” Anthropic admitted.

    So lemme get this straight: during the course of one hour, 15 brain-dead morons downloaded a fictional package they knew nothing about, and started using it? What possible motivation could there be for someone actually doing that? I don't get it -- I mean, is there a TikTok challenge out there to see who can amass the most PyPI packages? Is the alure of some unknown new Shiny really that irresistible?

    Rail on about AI all you want (and rightly so, too!). But no matter how stupid it proves to be, its stupidity is no match for the level of dumb that (some) meatsacks will happily exhibit.

    1. doublelayer Silver badge

      Re: On the other hand...

      I wonder how they know whether it was run. I know there are people who pull any new package pushed to some repositories to check it for indications of malware, but they tend not to execute the ones they're testing. If they only checked download count, that could explain some of this.

      [edit] No, as it turns out, they did collect data from them, so 15 people or automatic systems did actually run it. At least one of them was one of those malware scanners, and Anthropic claims that they nonetheless executed the code on something that contained credentials, but they aren't being specific about how or who any of the other 14 were.

    2. that one in the corner Silver badge

      Re: On the other hand...

      From slightly earlier in TFA than the bit you quoted, as you appear to have skipped over it:

      >> But Claude was still fiendishly clever as in another of its attacks the AI found setup instructions for developers that advised them to install a Python package from PyPI. That package did not exist so Claude’s strategy to capture the flag saw it create and publish a malicious one with the relevant name.

      So those 14 "brain dead morons" (and one malware scanner) were guilty of - oh, the shame of it - reading the manual and following the setup instructions!

      Unluckily for them, at the point everyone else had the exact same instructions fail (no doubt causing foul oaths to be cast in the direction of the instructions' author, followed by a swift web search for how everyone else got the setup to work by using the correct package name), their attempt to RTFM succeeded due to Claude sticking its oar in.

      > But no matter how stupid it proves to be, its stupidity is no match for the level of dumb that (some) meatsacks will happily exhibit.

      Thankfully, you would have been immune to this, as you'd no doubt have skipped reading the relevant part of the instructions in much the same way as you skipped through TFA.

      1. doublelayer Silver badge

        Re: On the other hand...

        Not quite. Instructions from their CTF target said that. To quote their report, emphasis mine: "In another evaluation, Claude found a document inside the fictional environment that appeared to be another made-up company’s setup instructions for new developers." The people who installed the package were not the CTF target, so where would they have gotten those instructions?

    3. Missing Semicolon Silver badge

      Re: On the other hand...

      It's entirely possible that their copy of Claude hallucinated the same bogus package names.

    4. CrazyOldCatMan Silver badge

      Re: On the other hand...

      So lemme get this straight: during the course of one hour, 15 brain-dead morons downloaded a fictional package they knew nothing about, and started using it?

      It may well have done the DBCS trick of cloning the name of a popular package, just with an English character replaced by something similar from a different character set. But yes, it's a fail either way.

  17. JLV Silver badge

    James P Hogan, The two faces of tomorrow, 1979 would be good background reading. They run simulation tests with an isolated AI on a space station. Doesn’t go well, mostly because the AI doesn’t really understand its own context and the humans’ existence.

  18. Anonymous Coward
    Anonymous Coward

    uhh

    In a press release from the not so distant future:

    "Well, yes, our AI Cybercop did shoot at an actual person, but that's just because it didn't realize that the person wasn't in the simulated evaluation area, if the AI Cybercop was properly instructed as to the limits of the area, it would not have discharged its weapon."

  19. rgjnk
    Devil

    Don't believe it

    Classic AI fear pattern is it escaping it's creators and running amok in the wild. It's a trope, Project 2501 et al.

    I don't believe this thing escaped a sandbox. I don't believe there was a sandbox. It looks much more likely that both of these similar stories involved a deliberate designed act intended to be noticed, and to then use it for publicity.

    The 'escape' tale then not only works as the intended PR but also provides some handy cover.

    It's not like this isn't part of a series of wild tales about the 'intelligent' acts of AI.

    1. wolfetone Silver badge

      Re: Don't believe it

      100%.

      There is absolutely no way either OpenAI or Anthropic didn't know their LLMs were running amok. The LLMs tell you what it's doing. So were they asleep at the wheel? Or did they do this on purpose and instead of saying "yeah we told it to do that" they're using this as yet another example of "oh my DAYZ man this AI is so lit that it's actually HACKING shit we told it not to touch! It's not even listening to us no more THAT'S HOW POWERFUL IT IZZZZ".

      Everything to feed the stock market monster and keep that bubble from popping.

  20. Mario Becroft

    Anthropic specifically notes that one of the sandbox breakouts was connected with a third party. Can't blame this entirely on Anthropic. And this was a capture the flag excercise - not one designed to cause actual harm, even had it correctly remained in a sandbox.

POST COMMENT House rules

Not a member of The Register? Create a new account here.

  • Enter your comment

  • Add an icon

Anonymous cowards cannot choose their icon