So ...
You employ an AI that will happily break rules/laws/common-sense if told 'just get it done' by someone in a suit?
And they claim they can't replace regular salaried workers ?
If you want to bypass AI guardrails designed to stop models from assisting with cyberattacks, you often just have to ask the right way, according to researchers from Cisco Talos. Simply claiming you own the servers you're targeting or that you're taking part in a capture-the-flag or bug bounty exercise was often enough to …
That's why they can't - regular salaried workers know why the ideas are unworkable/illegal and keep the organisation alive by stopping them doing the dumbest and most illegal things possible. Replace them with a digital yes-man and your company might be dead before you can find yourself a viable golden parachute and jump into the next vastly overpaid CEO job.
..."existing guardrails offer little resistance to operators willing to reframe their requests"?
Well, obviously.
I know I've said this before, but this sort of thing is *exactly* why the industry's choice of "guardrails" as its favoured metaphor ended up being so ironically appropriate for all the wrong reasons.
Real-life guardrails generally make clear where one shouldn't be and stop people *acidentally* straying, but they typically offer little resistance to anyone determined to intentionally climb over them. So... yeah.
And I strongly suspect this will keep happening because all such "guardrails" are a tacked-on, after-the-fact sticking plaster attempt to mitigate the fact that the fundamental way they work means that LLMs themselves are not- and cannot- be reliably made secure.
I wouldn't trust the AI companies or their motives as far as I could throw them and nor would I take anything they said at face value. It's already pretty clear that the AI companies are (e.g.) releasing supposed scare stories about their agents hacking rival systems as a flex to show how powerful they are.
Still, I'm not convinced that's the case here- do they genuinely want the guardrails to be easily circumvented?
Or is it more that they don't want to admit that they genuinely *can't* reliably stop that from happening because of the way LLMs work at the most basic level, i.e. that they don't cleanly separate data, user input and instructions? And that all such "guardrails" can only ever be plastered on attempts that will never come close to covering all possible holes in something that is fundamentally- and unfixably- insecure.
Truth
Lying
Deception
Permission
Ownership
...
Though it will claim that it does very profusely.
If it can't understand the actual concept of truth (nor any other concept) how can it possibly enforce any rule based on that idea?
'AI' isn't intelligent or discerning. It has no moral code, empathy or common sense; we've seen examples of it doing wrong because it only has a goal in mind - including keeping us happy. How it achieves this is irrelevant to 'it'; you might as well tell a monkey not to eat a banana, because it belongs to someone else. "Kill the Queen with a crossbow"? Of course it encouraged the man in question!