The Register Home Page

back to article Bypassing AI guardrails is so easy a script kiddie can do it

If you want to bypass AI guardrails designed to stop models from assisting with cyberattacks, you often just have to ask the right way, according to researchers from Cisco Talos. Simply claiming you own the servers you're targeting or that you're taking part in a capture-the-flag or bug bounty exercise was often enough to …

  1. Yet Another Anonymous coward Silver badge

    So ...

    You employ an AI that will happily break rules/laws/common-sense if told 'just get it done' by someone in a suit?

    And they claim they can't replace regular salaried workers ?

    1. breakfast Silver badge

      Re: So ...

      That's why they can't - regular salaried workers know why the ideas are unworkable/illegal and keep the organisation alive by stopping them doing the dumbest and most illegal things possible. Replace them with a digital yes-man and your company might be dead before you can find yourself a viable golden parachute and jump into the next vastly overpaid CEO job.

  2. ecofeco Silver badge
    Facepalm

    I, for one, welcome...

    ...our script kiddie overlords.

    I'm sure it will end well.

    1. PM.

      Re: I, for one, welcome...

      Certainly !

  3. Michael Strorm Silver badge

    "Bypassing AI guardrails is so easy a script kiddie can do it" and...

    ..."existing guardrails offer little resistance to operators willing to reframe their requests"?

    Well, obviously.

    I know I've said this before, but this sort of thing is *exactly* why the industry's choice of "guardrails" as its favoured metaphor ended up being so ironically appropriate for all the wrong reasons.

    Real-life guardrails generally make clear where one shouldn't be and stop people *acidentally* straying, but they typically offer little resistance to anyone determined to intentionally climb over them. So... yeah.

    And I strongly suspect this will keep happening because all such "guardrails" are a tacked-on, after-the-fact sticking plaster attempt to mitigate the fact that the fundamental way they work means that LLMs themselves are not- and cannot- be reliably made secure.

  4. Arkitekt

    AI guardrails are meant only to be a bare minimum effort to convince the public they are "trying"... they WANT them circumvented.

    1. Michael Strorm Silver badge

      I wouldn't trust the AI companies or their motives as far as I could throw them and nor would I take anything they said at face value. It's already pretty clear that the AI companies are (e.g.) releasing supposed scare stories about their agents hacking rival systems as a flex to show how powerful they are.

      Still, I'm not convinced that's the case here- do they genuinely want the guardrails to be easily circumvented?

      Or is it more that they don't want to admit that they genuinely *can't* reliably stop that from happening because of the way LLMs work at the most basic level, i.e. that they don't cleanly separate data, user input and instructions? And that all such "guardrails" can only ever be plastered on attempts that will never come close to covering all possible holes in something that is fundamentally- and unfixably- insecure.

  5. Pascal Monett Silver badge
    FAIL

    AI guardrails

    Those two words have nothing to do together in the same sentence.

    Given that AI companies have no idea how their bastard offspring works, the idea that they can impose guardrails is laughable and ridiculous.

  6. drguyrope

    Auto complete on steroids has no conceptual understanding of...

    Truth

    Lying

    Deception

    Permission

    Ownership

    ...

    Though it will claim that it does very profusely.

    If it can't understand the actual concept of truth (nor any other concept) how can it possibly enforce any rule based on that idea?

    1. Gavsky

      Re: Auto complete on steroids has no conceptual understanding of...

      But, but! AI has achieved singularity! Or, nearly - or some other complete BS...

  7. 47ufCapacitor

    The malware/info stealer chat groups I'm in bank on these bypasses to generate juicy .py scripts which do fun things

  8. Gavsky

    'AI' isn't intelligent or discerning. It has no moral code, empathy or common sense; we've seen examples of it doing wrong because it only has a goal in mind - including keeping us happy. How it achieves this is irrelevant to 'it'; you might as well tell a monkey not to eat a banana, because it belongs to someone else. "Kill the Queen with a crossbow"? Of course it encouraged the man in question!

POST COMMENT House rules

Not a member of The Register? Create a new account here.

  • Enter your comment

  • Add an icon

Anonymous cowards cannot choose their icon