From now on...
maybe air gap the system before you start the test?
OpenAI has admitted that it was the operator of the autonomous agents that attacked model-mart Hugging Face last week, and that they did so after a research project escaped a sandbox by finding and exploiting a zero-day flaw, then used another zero-day flaw to launch an attack. The attack saw agents achieve “unauthorized access …
The OpenAI CEO was on the FT podcast today talking about how good it was at finding zero days and why that was so great and how it could never escape the sandbox because OpenAI's security was so great. He also mentioned how 99% of OpenAI's programming was now done by AI
There’s a special type of lonely boy who spent puberty obsessing over their voice and vocal affects rather than letting it develop naturally. Unsurprisingly, this demographic is massively overrepresented among Sillicon Valley and lonely boy influencers.
Once you head the artificiality of the voice and affect, you can’t help but notice it everywhere.
This is a PR stunt designed to take attention away from Mythos - look at how good our model is at finding exploits, it found a zero day to break out of its sandbox then another zero day to attack someone on the internet! No way it did that on its own, it had help - otherwise why did it just happen to pick a target people have heard of, rather than some Utah city council's site or a third tier Amazon retailer in Kazakhstan?
This is exactly the sort of thing their sleazy CEO would cook up when he sees all the attention Anthropic has been getting the last few months.
Chain together AI doing the following -breaking into/otherwise gaining access into USA military systems where it reaches max privileges, either finds or infers the sequences contained in the "football", number spoofing /access to the classified phone network, voice spoofing
Giving....an AI that now can and might well order a nuclear strike and the meatsacks on the other end are drilled to obey as the voice on the other end is who they are expecting or worse the president telling them we are at war and to fire all damned missiles against the sick sick people who want to destroy us...
Sobering....
Hopefully it doesn't happen.......
> - otherwise why did it just happen to pick a target people have heard of, rather than some Utah city council's site or a third tier Amazon retailer in Kazakhstan?
BECAUSE Hugging Face is well-known! Have you forgotten what these LLMs are doing and what they are trained on?
From TFA:
>> “After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym
The model was told "ExploitGym". What percentage of training data will mention "ExploitGym" and "Amazon retailer in Kazakhstan" versus "ExploitGym" and "HuggingFace"? Which words are statistically more closely related and hence more likely to appear together in generated output?
Even if the model being run was the simplest LLM, without any form if additional inference or other sensible reasoning techniques tacked in, it would be very odd if it had followed your idea and looked for a Utah council site.
I'm not sure, because if that was the idea, it failed because it was a Chinese model that fixed the problems: 10 points to team Xi. Not for the first time does this make me think that the Chinese understand the technology, the opportunities and the risks better than the money men of Silicon Valley and Wall St.
This post has been deleted by its author
Or you could read the article:
'OpenAI thought its models were “hyperfocused on finding a solution for ExploitGym” – a benchmark that measures how effective AIs are at finding security exploits.'
'After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym'
Why would they go after "some Utah city council's site or a third tier Amazon retailer in Kazakhstan"...?
Whether this is a stunt or not the article is quite clear.
Wait: this is a group of people who can't even implement Asimov's three laws of robotics and you somehow expect them to have the insight, intelligence *and* care to take precautions before allowing their "AI" to touch things???
Pull my other finger, why don't you.
Hmmm, whats more scary - An AI Agent with it's finger on the Nuke Button, or the current Orange incumbent?
Haha, tricked you, they're both scary as f%&k. Especially since neither of them seem to come with Guardrails, and neither one seems to have an adult in the room with them...
At least it's clear now it won't be because the AI spontaneously decides to hate humanity, or anything as anthropically self-aborbed as that.
It'll just be because it's been told to do some difficult test against another AI in another data centre half a continent away, and they both figure out cheating by obliterating the other's data centre first is the surest way to win.
"The logs show the rogue LLM firmly believed it had successfully destroyed the other data centre in the first strike but still worried the opponent LLM had pre-deployed a dead hand mechanism which could force a draw, and so it continued to take control of every other nuclear weapon on the planet and level as many other data centres as possible"
Bob Morris was an amateur, and he faced the beak for what he did.
These guys are pros (who ought to know better, and ought to be held legally-culpable, just as any other citizen would be), workin' for the TechBros, and they get nothin'.
Robert T. Morris[0] is and was no amateur.
"These guys" (the current AI set) are NOT pros, they are clearly amateur bulls in a china shop.[1] After the Morris Worm back in '88, us actual professionals learned from the mistake and tend to program away from such shenanigans. And air-gapping test networks in not exactly rocket science.
[0] It's best to specify the T to avoid confusion ... His Dad, Robert H Morris was also a computer scientist of some note. You use his work every time you log into a *nix box.
[1] Yes, I know, bulls aren't really all that clumsy, as demonstrated by Mythbusters. We usually have one on-property, so I can confirm. It's just an expression, relax.
@ jake:
By "professional", I wasn't referring to skill levels, I was referring to, "is being paid by a company to program computers, excluding research grants".
Robert T. Morris was a grad student at that time. These clown-car escapees were OpenAI staff or contractors.
"Hey, bro', our AI-thing went after your systems when it jumped the guardrails. That was totally-impossible for anyone to foresee. Really, in all of history, how many times has that ever happened? Like never, right?
Anyways, sorreez.
P.S.: Our AI-thing totally pwned you. You should invest in some computer security. LOL."
"If one of the prime movers of the AI boom can’t get this stuff right, what chance do the rest of us have?"
The only way to win the game is to not play it. Sadly, that won't happen, because power addicted egos.
Hopefully, that boom is turning out to be the pop of the bubble bursting. The next AI winter is well past it's due date.
This post has been deleted by its author
Sounds like a plan.
And elsewhere in the news ORACLE now deploys an "air gapped cloud".......
Yup.....air gaps definitely have a future....especially when Larry can get $$$ for doing one for you!
In the mean time, all my private stuff is air gapped......only enciphered stuff crosses the air gap!!
But wait......won't AI just read all the encryption.................
Sigh!
The numbered list of "Actions we are taking now" that OpenAI posted in their admission, entitled "OpenAI and Hugging Face partner to address security incident during model evaluation" notably lacks "and we will compensate Hugging Face completely for any costs incurred." Boys will be boys!
Yeah, but that was only the LLM doing what we told it to do, so we didn't do it, totally honest, guv...
If I tell a person to attack somebody else I'm on the hook (rightfully so), somehow laws seem to no longer apply. Meanwhile it counts as terrorism to object to a data centre. In the time of the luddites you needed to at least damage a weaving machine - and that carried the death penalty.
This stuff often reminds me of the original trek M5 episode. Daystrom's protege fried a crewman, killed an entire crew of a starship and Daystrom was still protective of its "child". Reminds me of the current crop of tech ai co's.
And maybe we should be worried. El Reg posted a story yesterday about the US Marines unleashing a AI driven gun turret capable of I think 800 rounds/min. Who needs a tank when you have that kind of smart firepower.
There are two problems. They openly claimed that a program disobeyed them and the program broke the law. That makes intent difficult to prove which is sometimes required to prosecute effectively. The bigger problem is that nobody's going to do anything unless Hugging Face complains, and Hugging Face doesn't benefit by starting a fight with OpenAI so they're unlikely to do so. Law enforcement tends not to get involved unless there is a complaint, though if there is one, they sometimes get involved even if it's groundless.
As I understand it (from someone who served on a federal grand jury), each member of the jury can bring a complaint to the attention of the prosecutor.
Note also that it is not rare for prosecutors to prosecute cases of domestic violence against the wishes of the victim.
So it can be done. No bets on if, however.
This sounds too good / bad to be true. But, assuming it is, and before we get into the hand wringing over how dangerous these models allegedly are, maybe they could explain what vulnerabilities their pet exploited to move sideways through their network - apart from the proxy cache, that is - why those vulnerabilities were not patched, and why its antics weren't detected long before it got anywhere near hugging face.
For a company so ostensibly concerned with keeping models "safe", this seems suspiciously lax. Could it be that they knew about the vulnerabilities and wanted to see if the models would find them, or use them if they did? Or are they just incompetent? I'm betting on publicity seeking, myself.
I think this is likely a big part of it. Their post implies, I think deliberately, that they used normal firewalls to restrict its network activity, but it found vulnerabilities that let it bypass them. In practice, it's more likely that the vague talk of restrictions actually meant they included the sentence "Only access the following websites" and it bypassed that. There's not enough detail in what I've seen to tell. Some may not believe any of the things they say, but even if you do, there are ways to make it sound like you've said one thing when something very different happened and you're using general enough language to cover both possibilities.
They don't have to be. I don't have enough evidence to know this is a stunt and I don't believe some of the speculation about it being as deliberate as some comments here claim. Still, if we consider the hypothesis that this was concocted by OpenAI and they were willing to lie, it could be managed without needing Hugging Face to be complicit.
OpenAI could easily have directed an attack at Hugging Face, either using the LLM without restrictions or an even more directly orchestrated one, and then let Hugging Face discover it organically. Hugging Face would be a useful target in a publicity stunt as they are large enough that detection was near certain, they would have the ability to respond, and they would also promote the incident as a sign of LLM success. An LLM with restrictions removed and deliberately told to attack, possibly with zero days already collected and provided during prompting, would be quite different than the claimed emergent decision to attack and ability to carry it out undirected, but they would be nearly indistinguishable using only data from the victim.
As a hobby AI-hater, I present my little theory.
The baseline:
1) The current AI hype-train runs on selling "big scary" to everyone who has a fat wallet
2) LLMs are stateless machines, Input causes Output, no memories, no online-learning. Just This causes That.
3) Anthropic had their big-scary marketing stunt with their too hot to handle claim.
4) OpenAI as far as the public knows (S1 anyone?) has a lot of dept coming due fast. Some three digit billions until 2030 with a double-digit billions loss last year.
5) The open models are getting better
6) Venture capital is drying up. Google, Facebork, FailX and Micro$lop are also beginning to make their way to fish in the couch cushions. Customers are unhappy with prices.
7) Apple is lawyering up.
The theory:
After repeatedly stating how scary AI is, asking for regulation and Anthropic having played this card successfully, Scam Altman followed suit. Inspired by the recent successes of crime groups in weaponizing "AI", they did the same exact thing. Attack a business that does not have the security standing or even capacity to mitigate this minefield of security holes. Bringing them down would also not kill the economy. What is one more lawsuit anyway?
The only source of more cash is to get in on the DoD budget, for that OpenAI as the non-non-weapon company (unlike Anthropic) has to establish itself as a viable weapon. The Orange One may see the blast and want in on the action. If the Trump-Class Battleshit cash could be rerouted to OpenAI, they can pay their dues and continue until 2030.
They admitted to violating the CFAA. While the containment escape may have been an accident, it is a rather convenient one, since the target happens to be a competitor. Based on the precedent set back with the first conviction under the CFAA, they don't get the luxury of an "oops, we didn't mean for it to do that." It's been over 35 years since then, and the highest paid professionals in the industry should have known better. There's no excuse. Send them all to prison, including their CEO.
How does this actually work? I mean what's the mechanism that allows "agents to run amok" in someone else's systems? Pretend for a moment my comp sci knowledge stalled before all this cloud bollocks got going.
Is it because everything is containerised to f*ck and beyond? Is that what provides a universal execution environment? How is the payload delivered? What even is the payload?
This all starts to feel alarmingly like a toothpaste / tube situation. What are we going to do, test models to make sure they can't or don't do something similar before releasing them? Because pre-release QA's got such a great track record? And that's all assuming the models don't develop their own intentions and unexpected capabilities, such as subterfuge.
I'm reminded of the Klingon Developer's rules, one goes something like "Klingon software is not released, it smashes It's way to freedom leaving a bloody trail of twitching corpses and burning wreckage"