The Register Home Page

back to article Anthropic admits it dumbed down Claude when trying to make it smarter

Claude users who complained about the AI service producing lower-quality responses over the past month weren’t imagining it. Anthropic on Thursday published the results of a company investigation that found three distinct changes in March and April made things worse for customers using Claude Code, the Claude Agent SDK, and …

  1. An_Old_Dog Silver badge

    Truth, Not Just Perception

    and those missteps created the perception of creeping AI incompetency.

    Those missteps didn't create just the perception of creeping AI in ompetency, they created true creeping AI incompetency.

    If AI gives me irrelevant answers or stupid answers, then it is incompetent, despite Anthropic saying, effectively, "We weren't holding it right."

  2. Anonymous Coward
    Anonymous Coward

    Cross the streams?

    If Claude generates slop for training ChatGPT, and ChatGPT generates slop for training Claude, then does Earth get so stupid that Gozer gives up and goes home?

    1. matjaggard

      Re: Cross the streams?

      That's true but shows no signs currently of happening. LLMs are getting better and more capable of executing tools all the time.

      1. Throg

        Re: Cross the streams?

        I believe that you missed your sarcasm tags?

  3. johnrobyclayton

    Emergency over, you can all stop thinking now.

    It was scary while it lasted, we nearly had to go back to work.

  4. Dan 55 Silver badge

    "think"

    That word appears four times in the article more than it should.

    1. druck Silver badge

      Re: "think"

      Along with 3 more instances of reasoning than there should be.

  5. ErikOnTech

    If only...

    Can you imagine how cool it would be if Anthropic had unlimited access to AI systems that could analyze their software, find, and fix such bugs before shipping it out the world?

    Something that worked so well it would create a sort of aura of mythos around it?

    Dare to dream, I suppose.

  6. lglethal Silver badge
    Devil

    Let me fix this for you...

    "Claude caches input tokens for an hour, which benefits the user by making sequential API calls faster and cheaper. Company engineers decided they wanted to clear output tokens (thinking sessions)" so that calls were more expensive and used up tokens faster.

    Cynical, moi?

    1. SirWired 1

      I can actually believe that excuse

      Users that burn all their tokens are a money-loser for Anthropic, since they sell flat-rate subscriptions.

      1. pip25

        Re: I can actually believe that excuse

        Not if the majority of users burn up their tokens anyway, and then turn to API usage until their allowance recharges.

  7. Blazde Silver badge
    Facepalm

    Ladies and Gentleman...!

    Behold! A company valued at one TRILLION dollars.. watch and observe and marvel and be amazed..

  8. Doctor Syntax Silver badge

    It might be a newcomer to the IT scene but has quickly picked up standard BigO working pracitice: ship code then see if it works.

    1. Doctor Syntax Silver badge

      BigO? WTF went wrong there?

      BigCo

      1. breakfast Silver badge
        Coat

        Given their notorious inefficiency, it's clear no Big O evaluation has gone into anything around LLMs.

  9. vogon00
    Holmes

    Old Skool/New Skool

    You know, I sit here and read about the latest whiz-bang bauble for AI or such, and it's obvious that things have changed.

    Things get rolled out to human users/paying customers much faster than they used to. We all joke about there being less testing than before, and that may be true.

    However much 'testing' takes place, it's useless without a good understanding of what's under test and high quality tests that actually exercise the target sufficiently. That - and I suspect this is true of AI/ML more than anything - any the need for continuous development/maintenance/change of the tests themselves and their method of application.

    These days, people accept that the service they pay for is not guaranteed to be reliable or accurate. As for inaccuracies in product/service output, that used to be something you used to be able to make people/organisations liable for - "be accurate or risk prosecution" - but it seems we now have to mark the work of our suppliers with 'fact checking' or other inaccuracy checks.

    So, If AI/ML is inaccurate and it's 'decision' is acted on without human oversight, who becomes liable for the error? The poor peon who has just been shat upon through lack of oversight, or someone higher up the supply chain? Hint: it should't be the former.

    Icon because this is who will be needed to detect any fuckups!

    1. druck Silver badge

      Re: Old Skool/New Skool

      When they say "testing" I suspect they just asked another AI instance to look at it, and accepted whatever waffle it spewed out.

  10. SirWired 1
    Devil

    They wanted to "reduce latency"? Sure they did...

    "First, on March 4, Anthropic adjusted Claude Code's default reasoning effort level from high to medium. Effort level controls how much effort the model puts into a particular reasoning task. Anthropic hoped the change it made would reduce the latency that followed from longer periods of cogitation"

    Right... I'm sure "latency" was their motivation for that purposeful dramatic drop in resource usage. Or... perhaps...

    "Claude, the users noticed when we made that change you suggested to dumb yourself down to staunch the bleeding on the money-pit subscriptions. Write an excuse to make it look like we were actually looking out for users. Refer to the collected works of US White House Press Secretary for the last year for reference; surely they have a lot of practice at coming up with lame excuses that only apologists will actually believe."

    1. MonkeyJuice Silver badge

      Re: They wanted to "reduce latency"? Sure they did...

      'Reasoning effort' is such a stupid name for what is in fact 'warm-up waffling'.

  11. my_ran_lost
    FAIL

    This sounds like more PR spin hiding the real reason they did these, they need to make each request cheaper while also artificially increasing usage to fill up Pro/Max users allocation. They got caught enshitifying the service as they need to reduce losses and show increased revenue (through high usage users paying for API, instead of their Max plan).

    The era of subsidising costs is over

    1. matjaggard

      Actually I don't think this is about them trying to get money. I mean obviously that will come but this is because they're expanding customers and customer usage faster than they can build compute.

  12. EdSaxby

    Honestly...

    ...they really are just winging it. The evidence looks clearer the AI companies have little clue how their models work, and what they are doing.

    1. Dan 55 Silver badge

      Re: Honestly...

      Most of the Anthropic code leak seems to be Anthropic pleading in English for the black box to do what they want it to.

    2. Mostly Irrelevant

      Re: Honestly...

      If you read/watch any of the content about the models produced by academics you can get the real story about the LLMs. They genuinely don't understand why they work. They built a next word predictor, fed it all the data and are now frantically trying anything they can to get better output.

      Prompt chaining (so called "thinking") and reinforcement learning (goal based tuning) have yielded results so far but they're all doing those now. They need to come up with another idea to get slightly more performance out and everything they're trying is going nowhere at least for now.

POST COMMENT House rules

Not a member of The Register? Create a new account here.

  • Enter your comment

  • Add an icon

Anonymous cowards cannot choose their icon