- Anthropic told everyone Mythos was dangerous because it's proficiency with biologics and cyber security
- Anthropic didn't release Mythos like everything else. They released a neutered fable. They didn't get rid of Mythos
- Anthropic opens a lab in SF
There was always a quesiton of "will the labs stop releasing their models and start building around them instead?" Yes - they already have. Anthropic is a biologic and cyber security company, in addition to intelligience.
Personally I wonder if they've been holding back a lot. Opus 5.5 was a good release after a little stagnation. Open AI releases good models and everyone says Anthropic sucks and -- Oh would you look at that - a better model finally and all of a sudden.
with how much competition, benchmaxxing, and increased compute for inference, I can't imagine that they are intentionally nerfing their own public models for any reason other than that they can't figure out how to package it into something the public can use. The training data for all these frontier models goes well into 2026 at this point.
Yes, both labs already have monstrous internal teacher models they don't sell for inference, this is generally acknowledged. They cut releases for the public just to keep revenues growing, it's not their actual frontier.
Ok. Given this hypothesis, why is the software they release generally considered crappy by competitor standards, benchmarks, and open source standards?
Claude Code is an awful codebase, has leaked its own source code multiple times, and scores the worst on number of tokens burned vs pass rate percentages.
I mean, on the one hand the tool may be vibecoded crap - I don't know, haven't checked, taking it at your word.
On the other hand, Opus 5.5 cracked zero-shotting proper LCARS interfaces that near-perfectly adhere to the franchise "design language" even in tiny details, while simultaneously being 100% functional following my admonitions about Airbus cockpit design rules and nuclear reactor control room standards.
So yeah, why wouldn't I use it? It works spectacularly well.
Probably because bad code that you create initially without thinking that it's a core piece of your stack becomes depended on for its crappy behavior, and then you can't change much without breaking workflows.
I seem to recall Fred Brooks talking about that experience with OS/360 JCL (maybe just straight up in The Mythical Man Month?).
If agents really are superpowerful at programming tasks why not just have it rewrite the tool that the majority of your customers use and have it recreate the bugs? I mean presumably its the primary force behind the current version so what's the major cost there?
I imagine it comes down to economics.. there isn't much upside to fixing the last 20% of issues that the dumber faster models are missing.
The cost to serve, latency profile ,and internal demand for a maximal intelligence model would probably keep it pointed at harder and more valuable problems most of the time.
Either that or they are throwing spaghetti at the wall to see what sticks ahead of the IPO. After all, if solving all diseases is the “total addressable market” then that sure helps.
Even if your comment is taken at face value (which it shouldn’t), they aren’t the only stakeholders.
I presume you are instead alluding to drug companies not wanting to cure diseases, since they then have no market to sell drugs into. But even then the discussion is more nuanced.
> Either that or they are throwing spaghetti at the wall to see what sticks
Now we know what this is: <a href="https://commons.wikimedia.org/wiki/File:Claude_AI_symbol.svg" rel="nofollow">https://commons.wikimedia.org/wiki/File:Claude_AI_symbol.svg
Their cyber security seems to be a much more materially interesting (and likely profitable) business than "intelligence". Unfortunately they've created a sort of mutually-assured-destruction racket where they take payments from both "sides" of any secured boundary.
I'm not holding my breath for the biology side of things, but I suppose it's possible they find interesting things.
Maybe. I haven't heard anything impressive from the bio side of my social network, nor can I see anything from the computational side (including my own understanding). Thankfully it'll be pretty obvious if they find something interesting.
It's pretty frustrating to do cyber security work and not have access to the best models. OAI is a little more liberal here and I was able to get access to daybreak-blue but I have to use the lesser last-generation models. Essentially this gives a small number of orgs a huge advantage in the 'application layer' for that domain (including OAI or ANT themselves).
nonethewiser · · focus · HN ↗
- Anthropic told everyone Mythos was dangerous because it's proficiency with biologics and cyber security
- Anthropic didn't release Mythos like everything else. They released a neutered fable. They didn't get rid of Mythos
- Anthropic opens a lab in SF
There was always a quesiton of "will the labs stop releasing their models and start building around them instead?" Yes - they already have. Anthropic is a biologic and cyber security company, in addition to intelligience.
Personally I wonder if they've been holding back a lot. Opus 5.5 was a good release after a little stagnation. Open AI releases good models and everyone says Anthropic sucks and -- Oh would you look at that - a better model finally and all of a sudden.
WarmWash · · focus · HN ↗
Apparently OAI is already building GPT7 and GPT8
pvab3 · · focus · HN ↗
bentt · · focus · HN ↗
Just beat the current winner by enough to own the spotlight for a bit, then start prepping for the next go round.
_rutinerad · · focus · HN ↗
bpodgursky · · focus · HN ↗
sowhat1 · · focus · HN ↗
Claude Code is an awful codebase, has leaked its own source code multiple times, and scores the worst on number of tokens burned vs pass rate percentages.
Is anyone even using their Figma competitor?
addaon · · focus · HN ↗
owebmaster · · focus · HN ↗
Yes. CC started to create canvases without me asking. The mockups look good (it's just html+css), the tool is vibecoded crap
TeMPOraL · · focus · HN ↗
On the other hand, Opus 5.5 cracked zero-shotting proper LCARS interfaces that near-perfectly adhere to the franchise "design language" even in tiny details, while simultaneously being 100% functional following my admonitions about Airbus cockpit design rules and nuclear reactor control room standards.
So yeah, why wouldn't I use it? It works spectacularly well.
IanCal · · focus · HN ↗
monocasa · · focus · HN ↗
I seem to recall Fred Brooks talking about that experience with OS/360 JCL (maybe just straight up in The Mythical Man Month?).
CoolestBeans · · focus · HN ↗
XenophileJKO · · focus · HN ↗
The cost to serve, latency profile ,and internal demand for a maximal intelligence model would probably keep it pointed at harder and more valuable problems most of the time.
_aavaa_ · · focus · HN ↗
jbs789 · · focus · HN ↗
Time will tell.
downrightmike · · focus · HN ↗
ejj28 · · focus · HN ↗
jbs789 · · focus · HN ↗
I presume you are instead alluding to drug companies not wanting to cure diseases, since they then have no market to sell drugs into. But even then the discussion is more nuanced.
chrisjj · · focus · HN ↗
Now we know what this is: <a href="https://commons.wikimedia.org/wiki/File:Claude_AI_symbol.svg" rel="nofollow">https://commons.wikimedia.org/wiki/File:Claude_AI_symbol.svg
throwaway27448 · · focus · HN ↗
I'm not holding my breath for the biology side of things, but I suppose it's possible they find interesting things.
yesbabyyes · · focus · HN ↗
Well, that's the thing--perhaps we should be.
throwaway27448 · · focus · HN ↗
siliconc0w · · focus · HN ↗