The Register Home Page

back to article Anthropic to release Mythos-class models to the public

Anthropic has revealed its intention to one day release models that match the performance of its Mythos bug-finding AI to the public, once it can make them safe. In case you came in late, in early April Anthropic announced it had developed a model called Mythos that is so good at finding security vulnerabilities in programming …

  1. elsergiovolador Silver badge

    Security by obscurity

    The framing assumes attackers don't already have functional equivalents, which sits oddly with the published research on LLM-assisted vulnerability discovery and the perfectly observable gap between "best in class" and "good enough to be ruinous." Defenders locked out of Project Glasswing simply wait their turn.

    The interesting question, conspicuously unaddressed, is who currently benefits from those 6,202 unpatched high-or-critical bugs sitting in projects that "underpin much of the internet." Allied governments get a name-check in the same paragraph as the planned access expansion. "Safeguards" is doing rather a lot of work here as a euphemism for "managed disclosure with first refusal for friendly signals intelligence."

    75 patches out of 530 disclosed, blamed on maintainer capacity, has the convenient secondary property that the exploitation window is being set by the same parties choosing who gets early access. Security by obscurity, tilted firmly towards the offensive side of allied cyber programmes, and dressed in the vocabulary of responsible disclosure.

    1. Bebu sa Ware Silver badge
      Coat

      Re: Security by obscurity

      > a euphemism for "managed disclosure with first refusal for friendly signals intelligence."

      As another el Rego commentard wrote elsewhere Principled security holes — I guess it depends on whether you have any… principles that is… not that one would expect, in this game, one's principals to.

    2. Clausewitz4.1 Bronze badge
      Devil

      Re: Security by obscurity

      "underpin much of the internet."

      That’s overkill, as the curl developer stated.

    3. Anonymous Coward
      Anonymous Coward

      Re: Security by obscurity

      They have probably pilfered it - using AI / from an Anthropic GitHub repository. Or perhaps asked nicely and Claude spilled it’s guts.

  2. DarthProcrast

    The cosa nostra would be proud

    "Nice infrastructure you have here. It would be a shame if something were to happen to it."

    If you don't pay for our frontier security model, you're about to fall victim to our frontier vulnerability model.

    It's a protection racket, plain and simple.

    1. ecofeco Silver badge

      Re: The cosa nostra would be proud

      I've learned a new phrase this month that seems to cover quite a lot of the enshitification taking place these days.

      Promissory estoppel.

      It was a huge, HUGE mistake allowing software makers to have exemption from it. And here we are.

  3. Sproggit Silver badge

    The is the AI Model

    Been working with different models for a couple of years now. Mostly Chat-GPT through employer subscription, more recently Claude.

    One thing I've noticed has been a steady - but non-uniform - erosion in the amount of work I get done per session. This is partly explained by existing reporting that shows that e.g. Anthropic reduced the token budget for the Pro model, but offset it by giving Max5 users more than the supposed 2x allowance [for 5x the price] of Pro.

    What I've experienced is that as both professional and personal projects have leaned in to AI more - to get the job done, to address particularly thorny issues - the "spend rate" on tokens and allowances has drastically increased.

    Real-world example for you. A few weeks back I registered for a free Claude account and used it to diagnose a config error with a Postfix+Dovecot mail server. Dozens of file content reviews and edits. The total session - on a Sunday - lasted OVER FIVE HOURS. On the "Free" account. So I subscribed to Pro.

    Today my morning session had lasted about an hour when I got the "90% of session limit warning". Claude and I had just been regression-testing some code mods and found issues. So my next instruction was: "Stop all fix activity. Document findings in text file for next session."

    Off it went.

    Next thing I saw on my screen was the "You're at your session limit. Buy more tokens or upgrade to Max".

    10% of a session to put some notes in a file? Are you serious?

    The problem here is that the formula that Anthropic, OpenAI and others are using to "measure" user consumption is obscure and capricious.

    And of course, the models are intelligent enough to scan code libraries and work out how much of the change is model-driven and how much by devs - and once they see that they have become a critical dependency, "That's a nice little project you have there. Do you have a deadline to meet? It would be a shame if something were to happen to it..."

    I smell something rotten.

    I know this is laughable, but we are going to need to push for regulation, oversight and transparency. These guys are as bad as UK telecoms providers and their laughable, "Up to 63Mb/s" claims.

    The issue here isn't that they are definitely committing fraud. The issue is that if they WERE committing fraud, it would be next to impossible for anyone to find out.

    I'm not typically a fan of increasing regulation, but given the impact of AI in our lives, I think it's a justified move.

    1. ecofeco Silver badge

      Re: The is the AI Model

      Just assume they ARE committing fraud until proven otherwise.

      Just don't wait for Godot, er I mean the proof they aren't. And never forget, nobody owes ANY company ANY benefit of a doubt. Ever.

    2. Clausewitz4.1 Bronze badge
      Devil

      Re: The is the AI Model

      ” I smell something rotten.”

      Not only you. The AI business model isn’t profitable.

      They are trying to IPO it before people realize it.

      So investors will foot the bill instead of big tech.

    3. FIA Silver badge

      Re: The is the AI Model

      I smell something rotten.

      You seem to think there's some kind of conspiricy going on?

      There isn't.

      None of these companies are profitable. They were always going to jack up the prices.

      It's called a 'loss leader'.

      10% of a session to put some notes in a file? Are you serious?

      Surely that's much more verbose than code? Why would you expect that to take fewer tokens? What experience in developing an LLM that suggests you you this is unrealistic? or is it just a 'gut feeling'?

      I know this is laughable, but we are going to need to push for regulation, oversight and transparency. These guys are as bad as UK telecoms providers and their laughable, "Up to 63Mb/s" claims.

      What do you mean by this? Again... none of these companies are making any money? How do you think they're going to bankroll all this stuff?? What would the regulation do? Force a company to be unprofitable for ever?? I don't understand. The entire history of tech has been like this, VCs push vast amounts of cash into the new shiny with a vague promise of future returns.

      If you like the AI, you're going to be paying for the AI soon. If you're sensible you've factored this in already.

  4. amanfromMars 1 Silver badge

    Get with the ProgramMING .... Don't Battle against the Inevitable and Desirable.

    Project Glasswing ... in its second-to-last paragraph reveals the company’s next step will see it “… work with critical partners – including US and allied governments

    That is not a progressive forward step. So therefore a real dumb move with partners unfit for future Greater IntelAIgent Games purpose more than just likely to leave one prey to that which, and/or those who, know what is to happen and how to create and driver it.

    Things aint the way they used to be ...... and the future is never going to be similar and familiar to the present either ...... no matter how much one might wish it to be so.

    1. Clausewitz4.1 Bronze badge
      Devil

      Re: Get with the ProgramMING .... Don't Battle against the Inevitable and Desirable.

      Brazil is multipolar

  5. StewartWhite Silver badge
    FAIL

    Why are huge numbers of bugs in key systems acceptable?

    Why does nobody raise the point as to why we as an industry and those who buy systems think that companies shipping poor quality software is fine?

    This was always technical debt - it's now coming home to roost and everybody is clutching their pearls.

  6. Sproggit Silver badge

    Real world Example

    For anyone out there who isn't using AI for tasks such as coding support and might be a bit baffled by some of our complaints, I thought a real-world example [as best as I can manage] might help give an insight as to the obscurity and inexplicable "costs" when running sessions with a model.

    What you see below is an exchange at the end of an 90-minute session with Claude Opus, Anthropic's most powerful [and thus most token-expensive model].

    We've just spent > 80 minutes chasing a couple of defects to ground. One was in a multi-form, "Create New Record" type setup, written in a single re-entrant PHP script. The other was to diagnose a fault in a complex, context-sensitive and authorisation-driven menu framework (which I wrote and Claude broke). It was a very productive session and as you can see, all that problem diagnosis - and then a counted 7 cycles of me doing unit testing, failure-capture testing and regression testing on the "create" code, with Claude diagnosing reported oddities by scraping log files in real time.

    That was 74% of a session's token allowance.

    Then, as the log below shows, Claude:-

    1. Ran 7 commands [that's me giving Claude access to my BASH script library to reduce cost of doing stuff]

    2. Performed 3 git commits - that's literally entering 3 statements on my behalf and letting git do the heavy lifting...

    3. Documenting where we had reached and next steps by editing 4 files:-

    3.1. A summary of what we did this session

    3.2. Updates to our snagging list

    3.3. Updates to "CLAUDE.md", which is basically our "Big Rules for this Project", which I required Claude to read at every session start

    3.4. Putting down a skeleton of our "Next session brief" for me to edit/refine off-line.

    You're going to have to take my word for it that the above is trivial when compared with the 50+ minutes of deep triage and unit testing we just did.

    The entirety of what follows took maybe 2 minutes to complete.

    But... before I closed down the session, I went and checked my token usage - and I was apparently and miraculously at 100% - not a percent under, nor over. [You get an amber pop-up warning window when you actually hit 100%]. So that's 26% of a session's token allowance to create scratch notes of what we did and what comes next - basically a dump of what it's holding in session storage.

    I don't know for sure - I'm not sure how I could test this in real time - but it looks to me that the moment we hit my objectives for that session and the documentation was complete, token usage miraculously jumped to *exactly* 100%. Not 99%. Not 101%. *Exactly* 100%. What are the odds of that, eh? Twice now, in a week as a paying Claude user.

    I shall be checking every session from now on.

    Don't get me wrong, the Anthropic models are stunningly effective. But there are some deeply fishy behaviours when it comes to token consumption. Top tip for anyone starting out using Claude for this sort of thing: Take the time to read up on how to get the best out of your Model-of-Choice. Look for articles like this:-

    https://faizkhan3597.medium.com/98-of-your-claude-tokens-are-already-wasted-heres-how-to-take-them-back-6a758565d28c

    with plenty of concrete proposals, then experiment with them. There are quite a few YT videos out there with similar content. If you're paying for access, it's in your best interest to be aware of how to get the best of of the resource.

    Me: Re-test complete. urDefault is good; urCreatorRef is good.

    Me: Session Usage at 74%

    Claude: FP6 is signed off. Marking #9 and proceeding straight to the commit per the three-commit recommendation from the previous session's brief. Will go quickly to leave room for the end-of-session brief.

    Claude: Ran 6 commands, used 2 tools

    Claude: Ran 6 commands, used 2 tools

    Claude: Commit 1 — 25-May unblockers:

    Claude: Ran a command

    Claude: Ran a command

    Claude: Commit 2 — FP6 deliverable:

    Claude: Ran a command

    Claude: Ran a command

    Claude: Task list is in good shape. Continuing with commit 3 of 3.

    Claude: Ran 2 commands

    Claude: Ran 2 commands

    Claude: Three commits landed cleanly. Two backup files left untracked (your offline backups — safe to ignore or move to Archive).

    Claude: Used a tool

    Claude: Used a tool

    Claude: Now updating CLAUDE.md §4 with FP6 ✅ and new commit hashes, then writing the end-of-session brief.

    Claude: Edited 3 files

    Claude: Edited 3 files

    Claude: Now writing the end-of-session brief while we still have budget.

    Claude: Used 2 tools, created a file

    Claude: Used 2 tools, created a file

    Claude: All seven tasks complete. Session summary:

    Claude: FP6 ✅ closed and committed. Three commits:

    59f1f7d — 25-May menu/auth/dispatcher unblockers

    8224e59 — FP6 Add User multi-step form (with B3/B4/B5 fixes and the urDefault fix)

    d002816 — Migration 019

    Claude: Menu regression diagnosed (read-only, per your brief): both regressions have high-confidence root causes and small fix sketches in Daily Brief/2026-05-26 - Menu Regression Investigation.md. tracked as B-003 and B-004 in CLAUDE-snagging.md.

    Claude: Snagging framework created: CLAUDE-snagging.md with seven open defects (B-001…B-007) and one closed (B-000). Housekeeping cleaned up; CLAUDE.md companion list updated.

    Claude: Recommendation for next session: B-003 + B-004 fix as a quick standalone pair (one line + a six-line block removal). It removes the visual broken-window, restores the menu chain, and lets us test future modules via the navigation path rather than direct URLs — which would also reduce the practical impact of B-006 while FP7 dispatcher auth is being planned.

    Claude: End-of-session brief at Daily Brief/2026-05-26 - Session 1 - FP6 Close-out + Snagging Framework.md. CLAUDE.md §4 updated with the new commit hashes and FP6 ✅.

  7. Anonymous Coward
    Anonymous Coward

    "once it can make them safe". So, never then.

  8. eric.sherrill
    IT Angle

    If you can find it, can't you fix it?

    Why not prompt the AI to fix the bugs it finds? I know a flood of slop PRs doesn't help anybody, but if they could be reined in and validated like what they did for the vulns, it could help?

POST COMMENT House rules

Not a member of The Register? Create a new account here.

  • Enter your comment

  • Add an icon

Anonymous cowards cannot choose their icon