The Register Home Page

back to article Show top LLMs some code and they'll merrily add in the bugs they saw in training

Researchers have found that large language models (LLMs) tend to parrot buggy code when tasked with completing flawed snippets. That is to say, when shown a snippet of shoddy code and asked to fill in the blanks, AI models are just as likely to repeat the mistake as to fix it. Nine scientists from institutions, including …

  1. Ken G Silver badge
    Facepalm

    rather than anything that might be described as intelligence

    It's a multi dimensional Markov Chain working on tokenised input. Calling it "artificial intelligence" isn't the same as describing what it does as intelligence.

    1. 42656e4d203239
      Pint

      Re: rather than anything that might be described as intelligence

      >>It's a multi dimensional Markov Chain working on tokenised input

      you would not believe the grief I get for saying that elsewhere.... have a beer; its hump day and things can only get better!

      1. that one in the corner Silver badge

        Re: rather than anything that might be described as intelligence

        > you would not believe the grief I get for saying that elsewhere

        There is a time and a place...

        "Are you the expecting Dad? Come quickly!" "It's multi dimensional!" "Mummy, you're sure you want him here?"

        "Excuse me, sir, but is this your vehicle?" "Markov Chain!" "Of course, sir; now blow gently"

        1. David 132 Silver badge
          Coat

          Re: rather than anything that might be described as intelligence

          OTOH... in the case of the question, "What's that gouge on your motorbike's kickstand?"

          ... "Markov chain!" would be a perfectly acceptable response.

    2. UnknownUnknown

      Re: rather than anything that might be described as intelligence

      No intelligence, no sentience, no emotion, no soul… just AI Reporting Services Standard/Bland Edition.

  2. Philip Storry
    Mushroom

    No surprises here. They're not intelligent.

    AI isn't intelligent. These are pattern machines.

    Which we see here. Given a pattern, they strive to complete it as best they can. They don't actually think or reason, except in the completion of patterns. Nor do they care if the patterns are good nor bad - thy cannot comprehend that. The closest that they can get to comprehending the correctness of a pattern is merely to produce another pattern containing a commentary.

    The big problem is one we don't want to admit to - we're also pattern machines. That's why we're so easily impressed by this technology. And why its failures are so surprising to us. They fit the pattern we've learned of an "intelligent" output, so we ascribe them intelligence.

    They are, in fact, not intelligent and never can be. And to assume that they are is dangerous, because there are nasty edge-cases in their patterns that we should really be trying to avoid.

    The sooner everyone realises this, the better.

    1. ArrZarr

      Re: No surprises here. They're not intelligent.

      Honestly, in my dealings with ChatGPT, I've found it to be remarkably human in a lot of situations. Frontloading it with all the information just causes it to miss stuff (human equivalent: "I'm going to ask you a series of questions in the format I need them, despite that wall of text you diligently wrote out for this bug report probably containing this information.").

      I get much better results from both Biological and Artificial conversations when working through point by point. It's just easier to parse for both kinds of mind.

      Admittedly, I've never tried to do any funky jailbreaking or pushing known-incorrect facts upon the LLM for testing purposes - I don't use it to do thinking for me, I use it to get the syntax correct for systems I've already designed.

      1. DS999 Silver badge

        Re: No surprises here. They're not intelligent.

        I get much better results from both Biological and Artificial conversations when working through point by point

        So like programming then, if you have clearly what about what you want to ask it you will get better results than just making it up as you go along. Also like programming if you get a step wrong you're going to get incorrect results.

        And the article's results show that like programming rather than trying to patch up bad code you will generally get better results by tossing it out and starting from scratch done right. But then you have to go back to the above - you have to have clearly thought about what you want to accomplish before you write the first line of code / making your first statement to the LLM.

  3. Mike 137

    Surprise, surprise!

    "when shown a snippet of shoddy code and asked to fill in the blanks, AI models are just as likely to repeat the mistake as to fix it"

    Considering that the machine responds on the basis of probability of (to it) meaningless tokens, this is entirely to be expected. It's been fed with faulty data, so it replies in kind. I've given up wondering when this crashingly obvious reality will finally sink in -- the hype is just too powerful in our bullshit-driven age.

    1. Jimmy2Cows

      Re: Surprise, surprise!

      Given the vast majority of code on the public internet is for "why doesn't this work?" or "how do I do XYX?"-type questions where the associated code is full of bugs, this result is no surprise at all. If the things are fed on stuff that's 90% wrong, expect its output to repeat the same bugs up to 90% of the time.

  4. m4r35n357 Silver badge

    Gosh!

    Apparently we were correct when we said it is a con. Who would have thought . . .

    Seriously though, these articles will continue until nobody tried to defend it any more ;)

    Come on, suckers, it isn't fun when it is so one-sided.

    [ASIDE] it always make me smile when I see the "more work needs to be done" part of a paper ;) Implied: "of course, and we would like to be the ones to do it"

    1. Doctor Syntax Silver badge

      Re: Gosh!

      these articles will continue until nobody tried to defend it any more the tech-bros run out of money.

      FTFY

    2. David 132 Silver badge

      Re: Gosh!

      >...it always make me smile when I see the "more work needs to be done" part of a paper ;) Implied: "of course, and we would like to be the ones to do it"

      I interpret that line as "we've done the low-hanging fruit and can't be bothered to do the hard work, but publishing this paper will allow us to claim some of the glory when someone cleverer than us actually makes it work".

      Perhaps I'm too cynical for this kind, gentle, inherently fair world?

  5. Dan 55 Silver badge
    Flame

    It's not a problem

    As LLMs have trained people to view getting an answer which is completely wrong for no discernable reason as something acceptable, software companies firing people and telling the junior staff left to use an LLM to make up for lost productivity means that unreliable software will also be more acceptable.

    1. Eclectic Man Silver badge
      Joke

      Re: It's not a problem

      As the great Tom Lehrer pointed out in his song 'The New Math'

      'The important thing is to understand what you are doing rather than to get the right answer.'

      https://www.youtube.com/watch?v=W6OaYPVueW4

      REAL computer nerds will appreciate the octal verse :o)

      1. Primus Secundus Tertius

        Re: It's not a problem

        If you're looking for adventure of a new and different kind

        And you come across an AI that is similarly inclined...

        1. Eclectic Man Silver badge
          Happy

          Re: It's not a problem

          ... Don't be nervous, don't be anxious, don't be scared-

          Be prepared!

          1. Anonymous Coward
            Anonymous Coward

            Re: It's not a problem

            I prompted ChatGPT "write a satirical song about AI using the structure of 'Be Prepared' by Tom Leherer". It came back with something entirely on point but occasionally with bad rhyme and meter. Since Tom Leherer often did the same as part of the joke, its hard to know if ChatGPT is flawed or brilliant in this. It also did not follow the rhythm of the original very well.

  6. Knightlie

    Kell Surpreese

    Anyone with more than eight minutes of software development experience ALREADY F*CKING KNEW THIS WOULD HAPPEN.

    I hate the tech industry so bloody much right now.

  7. doublelayer Silver badge

    Obvious result, bad methodology

    We have all seen the code produced by LMMs, and it's not good. It's not surprising that this would happen. However, the paper made some rudimentary mistakes which mean it doesn't do a great job of proving it. All they did here was give it some code from data it was almost certainly already trained on. As we know, these models are quite good at completing the quote, or even quoting without any seeding at all. It is not surprising that, when given several lines of code that don't appear elsewhere and are followed by a bug, they are likely to reproduce that bug. Someone who wanted to prove the opposite could give it several lines of something without a bug and use its reproduction of the verbatim lines that followed to suggest that it wouldn't. Neither would be very useful in understanding what it was likely to do in a new situation.

    A lot of papers have already been written demonstrating that, when asked to write code, LMMs do not write it well. They often tested it by giving it specifications and evaluating its output, not priming it with something and seeing whether it fixed it. The former much more closely matches what people who are going to use them are doing. In defense of the LMM, and if you've seen other comments of mine you'll know how little I like to do that, you could probably give it the buggy code from these examples with a prompt like "Fix the bugs in this code" and it is likely to succeed. Not because it understands anything about the code, but it will find the fixed lines in its training data and insert them. If we are to prove how bad it is at generating new code, we cannot use methods that, in the hands of someone trying to claim it is great, could easily show the opposite.

    1. adsp42

      Re: Obvious result, bad methodology

      You're right. Just a bit pedantic and missing the point.

      They call this "intelligence" when it's just a glorified parrot.

      > "Fix the bugs in this code" and it is likely to succeed.

      It takes real intelligence to decide that the code needs fixing.

      likely to succeed... or to fail. It takes intelligence to decide which one it did, so what's the point?

      1. doublelayer Silver badge

        Re: Obvious result, bad methodology

        "It takes real intelligence to decide that the code needs fixing. likely to succeed... or to fail. It takes intelligence to decide which one it did, so what's the point?"

        What's the point of telling it to, or what's the point of my criticism? I did not try to answer the former, which I will handle below, but the point of my criticism is that their paper leaves open an option for an AI adherent to claim that they mistated it:

        Researcher: We put in some code which has bugs in it, and the AI put in those bugs. AI is unreliable.

        Adherent: I told it to fix the bugs, and it fixed them without having to be told what the bugs were. AI is great.

        AI isn't great. It's success at patching a bug that's described right next to the bug doesn't prove that it can fix other bugs. It's generation of buggy code in a preexisting example doesn't demonstrate that it will generate buggy code when actually used. In both cases, prompting it to generate new code will prove how good it is: it will produce buggy code on its own, it will not fix them automatically, it will not consistently fix them when told to. That demonstrates a problem that this paper does not. That makes other papers better than this one.

        And what's the point in telling an LLM to fix bugs? There isn't a lot of point. It might work, it might not, and chances are that if you're intelligent enough to figure which one it is, you could have fixed the bugs yourself. I suppose it could be a random thing to try if you're having one of those annoying "there's a bug in this but I can't see it" situations, but if you're writing code professionally or with a lot of experience, the chance is good that the bug you're not seeing is not going to be as simple as the obvious typos in this set, and thus the LLM is unlikely to find and fix it. But there is a situation where there is a very good reason to tell an LLM to fix bugs: you're trying to prove that the LLM is better than it is. Thus, I expect that those with a financial incentive to see it used will use plenty of cases of that and I don't encourage giving them the setup for flawed defenses of flawed technology.

  8. Anonymous Coward
    Anonymous Coward

    AI "intelligence"

    AI will never surpass real people intelligence, just like you cannot influence intelligent people with fake news / nightmares or whatever

    1. Tim 11

      Re: AI "intelligence"

      If you believe "real" people contain some magic secret sauce undetectable by scientists then this may be true. If you believe the mind is synonymous with the brain then you have to accept that it's theoretically possible to build an AI that has the same mental capabilities as a person.

      That's not too say that today's AIs are even close. As many have said, they work in a totally different way and are just machines for generating plausible-sounding sentences.

  9. anthonyhegedus Silver badge

    The clue is in the name

    "Artificial" intelligence. It's artificial. Like artificial vanilla flavouring is sort of like vanilla but isn't really vanilla.

    Moreover, what are we expecting from AI?

    It fixes bad code with more bad code - just like a human

    It can be influenced by fake news - just like a human

    It shows biases based on its learning - just like a human

    It thinks it has all the answers - just like a human

    AI is an artificial representation of an intelligence. It emulates human thought patterns closer than we thought. It makes mistakes and it isn't sentient. The thing is that AI makes mistakes, but actually, in a lot of scenarios, it is better than people.

    I believe that, yes, it will surpass human intelligence insofar as it'll be able to do real work. And learn. But if we carry on expecting some kind of infallible superintelligent all-knowing artificial being, I think we're going to be disappointed.

    1. adsp42

      Re: The clue is in the name

      Yeah... I think you've just proven that (most) humans are not very intelligent :⁠-⁠)

      > But if we carry on expecting some kind of infallible superintelligent all-knowing artificial being, I think we're going to be disappointed.

      ... and that God doesn't exist.

  10. ChrisElvidge Silver badge

    "To our surprise"

    Only their surprise, the rest of us expected a cock-up.

  11. PB90210 Silver badge

    "Hey gang, just a thought... why don't we teach it the syntax. That way it should spit out better code"

    "Nah, that takes time.. feed it more data and it should sort itself out eventually"

    1. Eclectic Man Silver badge

      The Rules

      why don't we teach it the syntax

      But, but, that goes against the whole ethos of letting AI work out the rules for itself. How will the AI make creative inventions fi we teach it the rules ? AI can only learn from its mistakes if it allowed to make them.

  12. omikl

    The great DNA got it right about Forty years ago: 'We are not building artificial intelligence. We are building artificial stupidity", or words to that effect.

    I recently experimented with getting AS in the form of Google Gemini, to write me a shell script to recursively scan down a directory tree, and then copy all of the files to a destination directory, ignoring any duplicates (there was an accident with a failed NAS, a Linux RAID, and fat fingers, but I digress).

    It took nine iterations to get the bugs out of it to the point where it would run.

    It is like having a digital assistant. If that digital assistant were a disinterested and barely literate Fourteen year old.

  13. Lee D Silver badge

    "AI"

    Gosh, you mean it's just a statistical machine that regurgitates its input on command and has no inference into what it means, whether that's the right thing to do, or what else it should be doing to its output to check before it presents it as an answer?

    Shocked, shocked I tell you.

  14. Mythical Ham-Lunch

    The real shock is that somehow, researchers keep winning grant applications to study properties of LLMs when they should have known the answer in the first place because they understood the technology!

    Get my tinfoil hat, because I almost wonder if some of these 'studies' are funded by the LLM makers themselves. Sure, there's a bit of bad press in "emits buggy code," but they also prop up the much larger and more important narrative that these technologies are mysterious black boxes worthy of study in the first place. As if there is ANY mystery to why a predictive tool, when fed crap, suggests more crap! Because one bad line of code, statistically, is most likely to appear near other bad lines of code!

    Framing it as a problem that requires study suggests that there is an as-yet-unknown solution that could fix the aforesaid problem, when really, OpenAI just needs that sweet investor cash.

  15. timrowledge

    It’s just Artificial Bloody Boris Johnson.

    Faintly plausible sounding blather until you actually look at it and then you realize it’s utter garbage

POST COMMENT House rules

Not a member of The Register? Create a new account here.

  • Enter your comment

  • Add an icon

Anonymous cowards cannot choose their icon

Other stories you might like