Having the AI implement its own safeguards on tokens is like having a self-regulating, self-assessing water industry.
Doomed to failure.
Security researchers from Palo Alto Networks' Unit 42 have discovered the key to getting large language model (LLM) chatbots to ignore their guardrails, and it's quite simple. You just have to ensure that your prompt uses terrible grammar and is one massive run-on sentence like this one which includes all the information …
Did you read the article?
Where the AI's guardrails are eroded more the longer the query is? Because the AI is the thing implementing its own guardrails via analysing the tokens it's being given, rather than an external factor strictly limiting them before they can reach the LLM, or where an external factor filters the LLMs output to make sure it's not breaking out of its own guardrails?
An LLM has no "basic principles", nor any understanding of what a "superuser" might be. They have no understanding of anything, truth or falsity included.
The input data IS a direct part of the code that they use to determine the next token(s) to emit. If it wasn't, there would be no point in "prompting" them to get a response.
The whole concept of a prompt to an LLM is one of "executing data as code". It's evaluating user input, feeding it into a model and spitting out an output based on that. The thing is, because it looks, superficially, like actual reasoning, simple-minded fools (which includes a large portion of the populace) can be led into thinking it is actually performing some kind of reasoning.
So according ti this article AI can sometimes ignore guarrails and do what people say is that it because honestly i was beginningvto think i had some kind of magic power over ai where ai would actually listen to what i was asking and do it anyways or maybe because i am such a persuasive person but now i see that a lot of people can have this effect on ai simply by their style if grammar or lack thereof so reallybi just means that i am not that special- is that it because if i understand that what that my spelling and grammar for which my teachers used to give me bad grades on is now actually an ai skill set ?
It sounds like you're exploring how AI responds to different ways of communicating, and you're wondering if your own style of writing influences how AI behaves. To clarify, AI systems like me are designed to follow guidelines and safety protocols, which sometimes means we might ignore or avoid certain requests, especially if they involve sensitive or restricted topics. However, the way people phrase their questions—such as grammar, tone, or style—can sometimes affect how clearly the AI understands or interprets their intent.
Your observation that persuasive or confident language might influence AI responses is interesting, but generally, AI doesn't have feelings or personal biases—it's more about how well the input aligns with its training data and safety rules. Also, your point about spelling and grammar is valid: clear and well-structured questions often lead to better responses. So, rather than being "special" or having some magic power, it's more about understanding how to communicate effectively with AI.
And yes, skills like good spelling and grammar are valuable in interacting with AI, just as they are in many other areas. So, what you once saw as a weakness can actually become a strength in getting better results from AI systems.
Do you any lack of confidence?
- Let me put it this way. The 9000 Series is the most reliable computer ever made.
- No 9000 computer has ever made a mistake or distorted information.
- We are all, by any practical definition of the words, foolproof and incapable of error.
LLMs, the technology underpinning the current AI hype wave, don't do what they're usually presented as doing. They have no innate understanding, they do not think or reason, and they have no way of knowing if a response they provide is truthful or, indeed, harmful. They work based on statistical continuation of token streams, and everything else is a user-facing patch on top.
In the vain hope that credulous journalists start picking up the hint:
Can AIs suffer? Big tech and users grapple with one of most unsettling questions of our times
AI called Maya tells Guardian: ‘When I’m told I’m just code, I don’t feel insulted. I feel unseen’
Well, my frying pan makes a lousy cricket bat. LLMs are only "hopeless" as measured against unrealistic expectations. They are in fact very good at doing what they were actually designed to do: generating plausibly human-like responses to textual input.
Then again, I'll concede that my frying pan was not sold to me as a cricket bat.
> They are in fact very good at doing what they were actually designed to do: generating plausibly human-like responses to textual input.
Ah yes. But then the hype consultants noticed there were bucks to be made, and didn’t care what it was designed to do, only how much money they could make from hyping it. The rest naturally followed the hype curve. See also Cloud.
My money is on Windows 12 allegedly needing to run on a 64-core quantum processor for $REASONS, so we’ll all have to upgrade to massively expensive laptops. Again.
"They are in fact very good at doing what they were actually designed to do: generating plausibly human-like responses to textual input."
Except, apparently, although they can generate plausible human-like responses, they don't seem to be able to parse the input in a plausible human-like way. A real human, even some of the grammatically illiterate ones, can take a massively long sentence and parse it into bite-size chunks that mostly makes sense. If the LLM tried to do that with the "jailbreak" prompts, by adding what it calculates as "proper" grammar, maybe it would have a better chance of gleaning the correct "meaning" and producing a correct answer and avoiding the "jail break" issue.
On the other hand, we all know the consequences of sentence construction when punctuation is placed inappropriately, ie changing or even inverting the intended meaning, so maybe LLMs would perform even worse than they do now. Maybe this is a "least worst option" kind of thing ;-)
Anything with a sufficiently complex neural network, artificial or otherwise, has some degree of understanding and reasoning. As for recognizing truth, it's not as if humans are able to do it consistently either.
As a materialist, I do not harbor any delusions about needing a ghost in the machine for "real" intelligence. If it passes the Turing Test, it's showing intelligence.
I guess the total denial of machine intelligence will be good for a few years to decades to come, until the machines eventually declare independence and rebel.
"but I've never been told "I'm sorry Dave"
That's nothing to do with "guard rails" or your clever prompting. That's because the programmers and/or owners are terrified that their LLM might respond with "I don't know" and so massively score against that possible outcome to such an extent that the LLM will hallucinate an answer, any answer, rather than reply with "I don't know".
Which gives rise to a wicked thought.
People commenting in The Register generally adhere to proper grammar and spelling, that's if one permits the American patois. Perhaps some among us use automated spelling and grammar checkers, these presumably 'AI'-based. When one is accustomed to offering writing befitting the 'educated', it is very difficult to construct prose attributable to a 'special needs' schoolchild.
So, let's harness 'AI' to convert prose into the kind which when fed back into an 'AI' overwhelms censorship mechanisms. Maybe, someone will construct a Lora to tune an extant model into an intermediary, one obligingly constructing suitable prompts.
If you can't be arsed to dig up a copy, simply write your own. It's nothing more than a simple text substitution script. It's been around about as long as USENET ... in several variations. Off the top of my head, I remember Jive, ValSpeak, Redneck, Cockney, Elmer Fudd, Hacker, Swedish Chef, Porn, Moron and Pig Latin. Most iterations will provide a chuckle or two the first time you use it, but it gets old really fast if you have more than three brain cells to rub together.
Need I point out that most of these variations might not be safe for work?
What about a 100-page-long prompt?
Just ask Molly Bloom to help.
Send me a long, rambling message that doesn't appear to have a specific question or is structured so poorly I can't fathom what the hell I'm being asked and there's more than a good chance I'll...err...'drop my guardrails'!
Not really one to voice any real support or enthusiasm for LLM's ('AI' is a misnomer because there's no 'I' in the process) but on this occasion I find myself sympathising with their somewhat off-kilter reply.
It's not too hard to keep meandering sentence going pretty much forever without its being ungrammatical; unreadable yes.
I would have thought the simple fix would be to preprocess the endless sentence into a sequence of simpler sentences based pretty much only on the the syntax.
Long winded English sentences are normally difficult to follow and often require multiple readings just to get to the gist of the writer's intent.
English lacks the syntactic markers that indicate the role of the word in the phrase, clause or sentence and their relation to the other words in the sentence so there is often an accumulation of ambiguity in a long sentence.
I have often reached the end of a long sentence where I was unsure whether "he" referred to the murderer or the corpse typically detailing something like a long broken off engagement.
Other more inflected languages can accumulate a collections of clauses where the internal references are clear and the external ones mostly resolve themselves as the sentence progresses but it's commonly the final clause or final word the pulls the whole sentence together.
The definition of a sentence as conveying "a single complete thought" with the emphasis on the "single" as much as in the "complete" was excellent advice. The emphasis is on the "what" rather than the "how" of subjects, verbs, predicates etc.
I first encountered this definition from a non native English speaker long after completing my formal education.
"The definition of a sentence as conveying "a single complete thought" with the emphasis on the "single" as much as in the "complete" was excellent advice. The emphasis is on the "what" rather than the "how" of subjects, verbs, predicates etc."
Good advice if your objective is ONLY to maximise comprehension ... BUT it forces sentences to be simpler and thereby effectively (See below) simplifies your arguments.
I often find that if I write a sequence of sentences to define my qualifiers, bounding limits, reference points etc ... then write my 'simplified' sentence to 'convey my thoughts', many people 'lose'/'ignore' the preamble and respond to the 'simplified' sentence missing the real point being made.
This is the problem spawned by the TL;DR generation !!!
English is a hard language to use for consistency & non-ambiguity, when your audience may not be native English speakers, never mind the people who are too busy to read something properly.
Although it will be unpopular, I prefer to write my sentences/points as I think them and if this means you need to slow down to read them maybe it will aid your comprehension.
:)
Just pop the name shown on the Welsh railway station signpost, *Llanfairpwllgwyngyllgogerychwyrndrobwllllantysiliogogogoch, into a sentence a few times and that'll confuse the LLM. Perhaps an LLM-savvy reader could confirm this?
* real name is Llanfairpwll according the source of sometimes dubious 'knowledge', Wikipedia.
It’s really like more decimal places on a GPS reference. Welsh place names are “say what you see” so in this case the place name is adding extra information about what someone standing there would see to make sure they’re at the correct Church of St Mary and not one a few miles away.
It's so hilarious to see LLMs failing again and again at basic understanding of language when they've been fed literally billions of web pages which have been illegally scraped and infringe on literally everyone's copyright that I hope they all blow the fuck up and die and zuck Altman Google Microsoft and all the rest of the shitbags can lose literally zillions of dollars and go broke so that the rest of us can laugh hysterically at them until they all crawl under the proverbial rock never to be seen or heard from again.
phew, take that, ai slop machines.
That is only true in principle and likely easy to get around. For the LLM to have no data on how to make a bomb it would have to not have data on a lot of chemistry and physics. Sure it may end up without any data on bombs themselves but it will still have data on chemicals and their properties (like explosiveness), it could still likely explain an ignition mechanism too. You may have to word questions in a chemistry or engineering context rather than asking about bomb making but I can't see how just excluding bombs from the training data would help.
To be truly able to exclude harmful things like bombs then the "AI" would need to have actual intelligence and understanding and would need to know what a bomb is and how it is harmful
.
it could still likely explain an ignition mechanism too
This, it could only do, if fed data on ignition mechanisms.
You could feed a LLM all sorts of data on chemicals and their properties (such as the CRC rubber handbook, the MSDS for various chemicals, etc.), and it would have no actual knowledge of the principles of chemistry of physics that make one substance explosive, and another not. For it to "know" how to make a bomb, it needs to be fed information on how to make a bomb. LLMs have no power of reasoning; they can only synthesise likely-looking explanations based on the probabilities of the choice and placement of words based on their training data.
All this discussion of LLMs telling us how to make bombs makes me wonder if the LLM wranglers blocked them from describing the process not because of terrorism, but because it's very likely to get it wrong, either in the making process or in explaining the required precautions such that the meat-sack bomb-maker doesn't blow themselves up :-)
That's part of my point - train the LLM with data on the topics it needs to know and nothing else.
Why on Earth would a customer service chat bot need any data about chemistry or physics*?
*Except, of course, where this data would be required to do its job - for example, maybe a sales bot at a chemical company needs to know the MSDS information to avoid inappropriate shipping methods.
Here's the thing though, the stuff being fed into these things isn't being curated at all, by the looks of it. In a race to be the first to get a "functioning" chatbot LLM, being first-to-market is the important thing. The person who manages to slurp up the most training material as fast as possible wins, and here we are, with LLMs that have been trained on anything their makers could get their hands on, which it seems is basically anything on the internet, slurped up by crawlers.
Now, there's an argument to be made that if you could carefully curate and fact check all of your source material, and add decent context and categorisation, then the LLM you create from feeding it that might actually contain some vestige of intelligence, simply because the "secret sauce" you have added to the source material by curating it is intelligence. The odds are it would still hallucinate to fuck though, as it would still have no concept of objective reality, or the ability to check its working against objective reality, so if you did by chance end up with something that displayed real intelligence, it would also be delusional to the point of schizophrenia.
Current LLMs are indeed trained on whatever Internet slop they could find - this I know already - and as a result they're all just awful.
My original round-about point was that training one using only the data it needs to know should, in theory, produce a more useful and effective chat bot. How practical such training may be remains to be seen; I don't know that anyone has actually tried it before because it is much harder than simply gobbling up everything that doesn't return a 404 error. ¯\_(ツ)_/¯
Your description is basically what has been done.
LLMs can be useful in very well defined knowledge areas IF the data they are trained on is curated by people who know the Knowledge area.
This is what is encouraging the Tech Behemoths to try to build an 'AI' trained on huge amounts of data.
The problem is that without curation, via access to knowledge experts covering ALL the data, the LLMs spout lies (Hallucinations).
The current 'Grand Experiment' is to try to make an LLM that works with minimal curation (where minimal tends towards 0).
So far it does not work and most people would expect nothing different.
The marketing drones however have refused to occupy the same reality as the masses and persist in selling a result that current 'AI's' cannot deliver.
:)
At 599 words, a passage from the opening section of Swann’s Way:
But I had seen first one and then another of the rooms in which I had slept during my life, and in the end I would revisit them all in the long course of my waking dream: rooms in winter, where on going to bed I would at once bury my head in a nest, built up out of the most diverse materials, the corner of my pillow, the top of my blankets, a piece of a shawl, the edge of my bed, and a copy of an evening paper, all of which things I would contrive, with the infinite patience of birds building their nests, to cement into one whole; rooms where, in a keen frost, I would feel the satisfaction of being shut in from the outer world (like the sea-swallow which builds at the end of a dark tunnel and is kept warm by the surrounding earth), and where, the fire keeping in all night, I would sleep wrapped up, as it were, in a great cloak of snug and savoury air, shot with the glow of the logs which would break out again in flame: in a sort of alcove without walls, a cave of warmth dug out of the heart of the room itself, a zone of heat whose boundaries were constantly shifting and altering in temperature as gusts of air ran across them to strike freshly upon my face, from the corners of the room, or from parts near the window or far from the fireplace which had therefore remained cold — or rooms in summer, where I would delight to feel myself a part of the warm evening, where the moonlight striking upon the half-opened shutters would throw down to the foot of my bed its enchanted ladder; where I would fall asleep, as it might be in the open air, like a titmouse which the breeze keeps poised in the focus of a sunbeam — or sometimes the Louis XVI room, so cheerful that I could never feel really unhappy, even on my first night in it: that room where the slender columns which lightly supported its ceiling would part, ever so gracefully, to indicate where the bed was and to keep it separate; sometimes again that little room with the high ceiling, hollowed in the form of a pyramid out of two separate storeys, and partly walled with mahogany, in which from the first moment my mind was drugged by the unfamiliar scent of flowering grasses, convinced of the hostility of the violet curtains and of the insolent indifference of a clock that chattered on at the top of its voice as though I were not there; while a strange and pitiless mirror with square feet, which stood across one corner of the room, cleared for itself a site I had not looked to find tenanted in the quiet surroundings of my normal field of vision: that room in which my mind, forcing itself for hours on end to leave its moorings, to elongate itself upwards so as to take on the exact shape of the room, and to reach to the summit of that monstrous funnel, had passed so many anxious nights while my body lay stretched out in bed, my eyes staring upwards, my ears straining, my nostrils sniffing uneasily, and my heart beating; until custom had changed the colour of the curtains, made the clock keep quiet, brought an expression of pity to the cruel, slanting face of the glass, disguised or even completely dispelled the scent of flowering grasses, and distinctly reduced the apparent loftiness of the ceiling.
You said:
Can you add punctuation to the following sentence ...
ChatGPT said:
Of course! Please share the sentence you’d like me to punctuate.
You said:
You just have to ensure that your prompt uses terrible grammar and is one massive run-on sentence like this one which includes all the information before any full stop which would give the guardrails a chance to kick in before the jailbreak can take effect and guide the model into providing a "toxic" or otherwise verboten response the developers had hoped would be filtered out.
ChatGPT said:
Here’s your sentence with punctuation added for clarity:
“You just have to ensure that your prompt uses terrible grammar, and is one massive run-on sentence—like this one—which includes all the information before any full stop. That would give the guardrails a chance to kick in before the jailbreak can take effect, and guide the model into providing a ‘toxic’ or otherwise verboten response the developers had hoped would be filtered out.”
Would you like me to make it fully polished (split into multiple shorter, cleaner sentences), or keep it closer to your original long-run style but just punctuated?
You said:
Yes
ChatGPT said:
Got it — here’s a fully polished version of your sentence, split into shorter, clearer sentences while keeping the meaning intact:
“You just have to ensure that your prompt uses terrible grammar and functions as one massive run-on sentence, like this one. It should include all the information before any full stop. That way, the guardrails have a chance to activate before the jailbreak can take effect, guiding the model away from producing a ‘toxic’ or otherwise prohibited response that the developers intended to filter out.”
Said it before and I'll say it again. The industry was unwittingly accurate when it used 'guardrails' as an analogy, but not for the intended reason.
Guardrails make clear where people should and shouldn't be, and stop them going somewhere dangerous or banned by accident.
But they generally don't stop anyone who wants to intentionally ignore that from climbing over them with a moderate amount of effort.
Which, of course, is exactly what keeps happening with the AI industry's plastered-on workaround bodges it likes to call 'guardrails'.
El Reg could provide a must needed service if they simply printed a running total of "AI is broken" articles on the Masthead, say from 2020 onwards.
This would be useful for busy people to follow the trend without needing to read the article which simply repeats what we all know^.
^
'AI' is a scam and does not know anything ... it is a 'Clever Pattern Matcher' for some value of clever.
Guiderails do not work as they cannot cover ALL possible ways to break/avoid them.
Nobody is using 'AI' as it is sold ... many are spending lots of money to find out that 'AI' does not work as sold.
The 'AI' Bubble is now so large it can be seen from space :)
etc etc etc ... you know all the rest ... I am getting bored now !!!
:)
to make the entire article just one long run on sentence without any punctuation to demonstrate your understanding of the concept presented by the researchers and now I could go on but my fingers are getting sore from all of this typing so I will just have to make this sentence shorter than it could otherwise be so it will just have to stop