Claude Opus 5.5
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
Claude Opus 5.5
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
throwaway2027 · · focus · HN ↗
handfuloflight · · focus · HN ↗
cronin101 · · focus · HN ↗
staticman2 · · focus · HN ↗
rich_sasha · · focus · HN ↗
cmrdporcupine · · focus · HN ↗
danw1979 · · focus · HN ↗
Retr0id · · focus · HN ↗
aoeusnth1 · · focus · HN ↗
fghorow · · focus · HN ↗
esafak · · focus · HN ↗
fghorow · · focus · HN ↗
esafak · · focus · HN ↗
outworlder · · focus · HN ↗
ThouYS · · focus · HN ↗
danw1979 · · focus · HN ↗
hmokiguess · · focus · HN ↗
RGS1811 · · focus · HN ↗
carlos-menezes · · focus · HN ↗
sailfast · · focus · HN ↗
lgessler · · focus · HN ↗
The user is right. The outage is a real concern, and the issue is worse than we realized. Requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5 encountered elevated error rates. Worth stating plainly: these are not just models — they are load bearing rungs on the software development tooling ladder, and a blocker on this level makes the outage really bite.
One decision that is yours to make, not mine: should an email be drafted to Anthropic support? This issue has teeth, and a canonical handoff can land us where the main gate is no longer breaking silently.
jaapz · · focus · HN ↗
lgessler · · focus · HN ↗
aenis · · focus · HN ↗
loopmonster · · focus · HN ↗
bibimsz · · focus · HN ↗
tda · · focus · HN ↗
[dead]
Retr0id · · focus · HN ↗
variety8675 · · focus · HN ↗
gavinray · · focus · HN ↗
[dead]
emadabdulrahim · · focus · HN ↗
akhilome · · focus · HN ↗
Hopefully the output from vanilla 5.5 is as good as they claim. I’ll try out later tonight.
[1] <a href="https://kizi.to/claude-talks-too-much/" rel="nofollow">https://kizi.to/claude-talks-too-much/
m4tthumphrey · · focus · HN ↗
gruez · · focus · HN ↗
It's just a standard hero image + text for me, with no scrolling effects.
edit: @iAMkenough figured it out, it was because I have prefers-reduced-motion enabled.
thejazzman · · focus · HN ↗
EricBurnett · · focus · HN ↗
KyleTheDev · · focus · HN ↗
I agree that it's sort of stupid, not a fan.
ealready_value · · focus · HN ↗
giancarlostoro · · focus · HN ↗
iAMkenough · · focus · HN ↗
Everyone that doesn't gets served some animated bullshit.
gruez · · focus · HN ↗
Yep, you're right. I tried on my phone and got the scroll through image.
mbreese · · focus · HN ↗
For a marketing page, it’s not the worst UX I’ve seen, but still slightly annoying.
dionian · · focus · HN ↗
iAMkenough · · focus · HN ↗
thebitguru · · focus · HN ↗
halyconWays · · focus · HN ↗
josefresco · · focus · HN ↗
halyconWays · · focus · HN ↗
oefrha · · focus · HN ↗
swader999 · · focus · HN ↗
amluto · · focus · HN ↗
bibimsz · · focus · HN ↗
serchinastico · · focus · HN ↗
amluto · · focus · HN ↗
anon373839 · · focus · HN ↗
dbbk · · focus · HN ↗
re-thc · · focus · HN ↗
nozzlegear · · focus · HN ↗
re-thc · · focus · HN ↗
nozzlegear · · focus · HN ↗
petesergeant · · focus · HN ↗
Catloafdev · · focus · HN ↗
Sounds like they noticed the complaints. I'm curious to see what LLM-isms this one may have.
gekoxyz · · focus · HN ↗
ithkuil · · focus · HN ↗
mavamaarten · · focus · HN ↗
ithkuil · · focus · HN ↗
works quite well
brandon272 · · focus · HN ↗
qurren · · focus · HN ↗
gwking · · focus · HN ↗
I don't mean to pick on this comment in particular. The majority of my work day is now spent reading AI generated text, and I look at HN (too much!) because I want to read human commentary. Humans pretending to be obnoxious AI on repeat is net negative to say the least.
brandon272 · · focus · HN ↗
fragmede · · focus · HN ↗
drnick1 · · focus · HN ↗
username_my1 · · focus · HN ↗
and it's not about the verboseness (even though it obviously contributes to the fatigue and loss of focus), I swear the vocabulary of the llms change working on the same task on the same codebase significantly.
I wonder if there are studies around this.
ygouzerh · · focus · HN ↗
meric_ · · focus · HN ↗
<a href="https://openai.com/index/where-the-goblins-came-from/" rel="nofollow">https://openai.com/index/where-the-goblins-came-from/
Small quirks can quickly add up in posttraining if not caught. Although TBH with how obvious Claude language is, I do feel like this is something Anthropic probably noticed and just assumed people would not care about. Now that people have obviously cared, they're probably actively looking to alleviate it
Eliezer · · focus · HN ↗
j_heffe · · focus · HN ↗
adastra22 · · focus · HN ↗
aray07 · · focus · HN ↗
I wouldn’t be surprised if Opus 5 was trained on content written by other LLMs
dgroshev · · focus · HN ↗
> The Vercel target is hard-coded. That's common and not wrong, but it's opaque; nobody reading this later will know which Vercel project it belongs to, and if the project is recreated the target changes silently. A comment or a named variable would help.
> Pointing a DNS name at Vercel is only half the job. The domain also has to be added to the project in Vercel's dashboard, otherwise requests will arrive and Vercel will reject them. That step lives outside this code, so it's easy to forget.
> Finally, [CENSORED] existing only in production is slightly odd on the face of it. It may be perfectly deliberate (perhaps a single shared testing tool that only needs one public address), but if you're reviewing this rather than just reading it, that's worth confirming.
It has the same annoying cadence and writing style with slightly less prominent claudisms.
sashank_1509 · · focus · HN ↗
dgroshev · · focus · HN ↗
* Consider leaving a comment about the hard-coded Vercel target. It's not clear where does it come from.
* [This is just a bullshit point, because the domain is not "added to" Vercel, it's provided by Vercel]
* Are you sure that [CENSORED] is prod-only? The name suggests otherwise. [also, what "if you're reviewing this rather than just reading it" even means?]
wren6991 · · focus · HN ↗
It means "I'm treating you as lay-person punter, not a developer working on this project." Opus 5 feels like it's constantly trying to reward-hack me into treating it as intellectually honest and epistemically humble, while in the same breath it talks down to you and smuggles its own bullshit assumptions and assertions into the conversation unchallenged. No progress on this front apparently. Glad I cancelled.
dgroshev · · focus · HN ↗
Claude is just comically bad nowadays.
redox99 · · focus · HN ↗
cruffle_duffle · · focus · HN ↗
Seems like it based on my first session. It still does the whole “bury the important thing in a pile of words” coupled with the “it might actually be important” thing… so basically you never really know what it’s talking about.
Honestly I trust opus so little that the entire “opus” brand is completely tarnished. Its writing style is so god awful that it needs more than just a point release. Either dump the name and ship a different model entirely or at minimum call it “opus 6”. Calling it 5.5 makes it sound like it’s basically a continuation of the same garbage output that 5.1 had but with some minor adjustments. And based on my single first test, that is what it appears like to me.
Gander5739 · · focus · HN ↗
tomhow · · focus · HN ↗
prodigycorp · · focus · HN ↗
km144 · · focus · HN ↗
tomhow · · focus · HN ↗
Lord_Zero · · focus · HN ↗
wtarreau · · focus · HN ↗
alpineman · · focus · HN ↗
skunkworker · · focus · HN ↗
Is the Xbox 360 (Xbox 2) vs PS3 debacle all over again.
ekckekcjekfj · · focus · HN ↗
It was odd at the time, yes, but no one really minded it truly. Heck, Xbox ONE was more of a fiasco/debacle than 360.
I see what you’re trying to get at with this comparison, but a “debacle” it ain’t.
meerita · · focus · HN ↗
meerita · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
buntp · · focus · HN ↗
wren6991 · · focus · HN ↗
FergusArgyll · · focus · HN ↗
frshgts · · focus · HN ↗
glub · · focus · HN ↗
joshstrange · · focus · HN ↗
> Input and output tokens are $4 and $20 per million, 20% less than Opus 5. Cache reads (which make up the majority of agentic and coding work costs) are $0.20 per million tokens, 60% less than Opus 5. Opus 5.5 also generates output more than 30% faster than Opus 5.
Better than Fable, cheaper than even the last Opus. I use Opus as my main driver so this is very exciting!
bayesianbot · · focus · HN ↗
bleonard · · focus · HN ↗
So longer threads get cheaper and one-shots stay the same price.
sznio · · focus · HN ↗
system2 · · focus · HN ↗
lanyard-textile · · focus · HN ↗
ygouzerh · · focus · HN ↗
mavamaarten · · focus · HN ↗
Zambyte · · focus · HN ↗
adastra22 · · focus · HN ↗
sznio · · focus · HN ↗
booty · · focus · HN ↗
copperx · · focus · HN ↗
NorwegianDude · · focus · HN ↗
The open models are getting closer and closer, and because they're open, people are not forced to pay the silly markup that is often over 1000x the cost to serve the model.
enraged_camel · · focus · HN ↗
anthonypasq · · focus · HN ↗
enraged_camel · · focus · HN ↗
Game_Ender · · focus · HN ↗
copperx · · focus · HN ↗
mgw · · focus · HN ↗
Maybe Anthropic finally felt the pressure from MiMo, DeepSeek, GLM Flash and Luna.
kibae · · focus · HN ↗
mudkipdev · · focus · HN ↗
simianwords · · focus · HN ↗
enraged_camel · · focus · HN ↗
sharkjacobs · · focus · HN ↗
God I hope so
abtinf · · focus · HN ↗
I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.
I don’t know how much better a model would have to be to get me to move off OpenAI at this point, but doing just a little bit better in terminal bench 4 isn’t it. It would have to be a difference in kind, or opening up the harness restrictions, or privacy guarantees (comparable to offline models).
sharkjacobs · · focus · HN ↗
God I hope so
kantahayashi · · focus · HN ↗
bushido · · focus · HN ↗
<a href="https://github.com/AminBlg/SimpleEnglish" rel="nofollow">https://github.com/AminBlg/SimpleEnglish
bkishan · · focus · HN ↗
rfgplk · · focus · HN ↗
RugnirViking · · focus · HN ↗
mavamaarten · · focus · HN ↗
lgessler · · focus · HN ↗
bibimsz · · focus · HN ↗
nonethewiser · · focus · HN ↗
boc · · focus · HN ↗
fastball · · focus · HN ↗
mikeocool · · focus · HN ↗
neilellis · · focus · HN ↗
drbscl · · focus · HN ↗
Trasmatta · · focus · HN ↗
unddoch · · focus · HN ↗
thefourthchime · · focus · HN ↗
abtinf · · focus · HN ↗
I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.
I don’t know how much better a model would have to be to get me to move off OpenAI at this point, but doing just a little bit better in terminal bench 4 isn’t it. It would have to be a difference in kind, like opening up the harness restrictions, or privacy guarantees (comparable to offline models).
Edit to address questions below:
ChatGPT supports oauth login.
Exe.dev has it built in. IIRC, pi also has it built in via /login.
felixgallo · · focus · HN ↗
ryanscio · · focus · HN ↗
felixgallo · · focus · HN ↗
Terminal-Bench 4.0 - Stanford & Laude Institute (with funding from all of the AI companies)
FrontierCode v1.1 - Cognition
CursorBench - Cursor (now SolarBoringSpaceXAI I believe)
GDPVal-AA - Artificial Analysis
AutomationBench - Zapier
Humanity's Last Exam - CAIS and Scale AI
Terminal-Bench-Science - Stanford, Laude, Ai2, Allen Institute
OSWOrld - XLANG Lab @ the University of Hong Kong
Chartography - Surge AI
esafak · · focus · HN ↗
abtinf · · focus · HN ↗
mattz56 · · focus · HN ↗
onlyrealcuzzo · · focus · HN ↗
This is news to me. Excited to try it out! Thanks.
nchmy · · focus · HN ↗
KeplerBoy · · focus · HN ↗
polalavik · · focus · HN ↗
abtinf · · focus · HN ↗
roughly · · focus · HN ↗
Can you give more details here? This sounds intriguing.
sidrag22 · · focus · HN ↗
So in simple terms, OpenAI doesn't restrict you to Codex, and gives their blessing to try whatever you want with their models(besides serving others with your subscription usage, that is still afaik against tos).
abtinf · · focus · HN ↗
Exe.dev has it built in. IIRC, pi also has it built in via /login.
cbg0 · · focus · HN ↗
qlte · · focus · HN ↗
copperx · · focus · HN ↗
That's an incredibly bold assumption.
cbg0 · · focus · HN ↗
qlte · · focus · HN ↗
<a href="https://artificialanalysis.ai/models/releases/claude-opus-5-5" rel="nofollow">https://artificialanalysis.ai/models/releases/claude-opus-5-...
And that assumes Opus 5.5 Medium is actually equivalent to Astra High in all real-world usage/personal work loads, which isn't guaranteed as benchmarks saturate. The High vs. High comparison (probably not equivalent, but for reference): If Opus 5.5 Medium isn't equal/better for what you're working on vs. Astra High across the board, the price difference would narrow a bit more each time you had to switch to High.So, if you're happy with Codex already it's not like Opus is now 1/2 the price and you'd be leaving a crazy amount of money/tokens on the table. Plus you have way more flexibility on the low end of the intelligence curve with GPT 5.6 Luna: Haiku (and Sonnet) can't touch that price/value ratio.
TuxSH · · focus · HN ↗
Gareth321 · · focus · HN ↗
Directly below this in the Cost per Intelligence Index Task table, the most efficient by far is Opus 5.5 Low.
margorczynski · · focus · HN ↗
jasbury · · focus · HN ↗
notatoad · · focus · HN ↗
abtinf · · focus · HN ↗
The Claude lock-in simply disqualifies anthropic entirely (for my use).
mlcruz · · focus · HN ↗
nonethewiser · · focus · HN ↗
>I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.
This is my experience with Claude code on my local machine. I suppose maybe you are doing something that naturally has system side effects? Obviously sandboxes have advantages sometimes but I havent seen a need for what I'm building.
abtinf · · focus · HN ↗
FWIW, the $20/month subscription also includes $20/month of LLM credits. That’s obviously not sustainable, but it should make it easier to try out the service. I would stick with them even if they dropped it.
Here is an invite link for a 30 day trial (that benefits me too if you were to become a paying member):
<a href="https://exe.dev/i/rlDF6GI5PGBZV4P" rel="nofollow">https://exe.dev/i/rlDF6GI5PGBZV4P
Or
ssh rlDF6GI5PGBZV4P@exe.dev
solenoid0937 · · focus · HN ↗
abtinf · · focus · HN ↗
GodelNumbering · · focus · HN ↗
If you are forced to reduce price despite raising capabilities, that certainly tells something about the market, and potentially about Anthropic future profitability too, since this model is their biggest topline contributor
alvis · · focus · HN ↗
liudaisuda · · focus · HN ↗
weiran · · focus · HN ↗
re-thc · · focus · HN ↗
For long running tasks it is. That's what made Deepseek so cheap.
vardalab · · focus · HN ↗
Espressosaurus · · focus · HN ↗
hedgehog · · focus · HN ↗
bayesianbot · · focus · HN ↗
blfr · · focus · HN ↗
rapfaria · · focus · HN ↗
If 5.5 is any better, I might try to do agentic-assisted development instead of just telling fable to delegate
blfr · · focus · HN ↗
neuronexmachina · · focus · HN ↗
herpdyderp · · focus · HN ↗
ascorbic · · focus · HN ↗
girvo · · focus · HN ↗
zanderwohl · · focus · HN ↗
locknitpicker · · focus · HN ↗
There are old wives' tales on how the original Fable was superb and the stuff of legend,but it as it was leaps and bounds beyond what other models were being offered then Anthropic opted replace it with a neutered version under the same name.
So today everyone can pay to use Fable, but legend has it they are paying for a nerfed replacement released under the same name.
coffeebeqn · · focus · HN ↗
btown · · focus · HN ↗
AJ007 · · focus · HN ↗
mcintyre1994 · · focus · HN ↗
drbscl · · focus · HN ↗
It does work out to be a similar cost per task though
jsnell · · focus · HN ↗
<a href="https://artificialanalysis.ai/models/claude-opus-5-5#intelligence-comparisons" rel="nofollow">https://artificialanalysis.ai/models/claude-opus-5-5#intelli...
It is most of the pareto frontier.
drbscl · · focus · HN ↗
93po · · focus · HN ↗
persedes · · focus · HN ↗
drbscl · · focus · HN ↗
persedes · · focus · HN ↗
naasking · · focus · HN ↗
<a href="https://artificialanalysis.ai/models/claude-opus-5-5?models=claude-opus-5-5-high%2Cclaude-opus-5-high#token-use" rel="nofollow">https://artificialanalysis.ai/models/claude-opus-5-5?models=...
piotrdz · · focus · HN ↗
johnbellone · · focus · HN ↗
epolanski · · focus · HN ↗
piotrdz · · focus · HN ↗
piotrdz · · focus · HN ↗
nl · · focus · HN ↗
The biggest proportional difference seems to be at max (5.5 is 38% more) and at high (5.5 is 21% less).
I think most people run at high and xhigh. At xhigh it is close enough to be task dependent and I don't think most people will notice. At high effort I think it looks like it will be an improvement for most people.
5.5 Max should probably be compared to Fable - it performs a lot better than 5 Max.
<a href="https://artificialanalysis.ai/models/claude-opus-5-5?models=claude-opus-5-5%2Cclaude-opus-5-5-xhigh%2Cclaude-opus-5-5-high%2Cclaude-opus-5%2Cclaude-opus-5-xhigh%2Cclaude-opus-5-high%2Cclaude-opus-5-5-medium%2Cclaude-opus-5-5-low%2Cclaude-opus-5-medium%2Cclaude-opus-5-low#intelligence-index-token-use-tabs" rel="nofollow">https://artificialanalysis.ai/models/claude-opus-5-5?models=...
make3 · · focus · HN ↗
meerita · · focus · HN ↗
gwd · · focus · HN ↗
Opus 5.5: Found 8/14 issues. Total cost: $15.40
Fable 5.1: Found 7/14 issues. Total cost: $66.34
Opus 5: Found 6/14 issues. Total cost: $15.19
Sonnet 5: Found 2/14 issues. Total cost: $19.15
This is a relatively small sample size, but it was both the best and the cheapest.
ETA: NB this is "Equivalent API" cost as reported by claude's CLI; I was using my subscription.
retinaros · · focus · HN ↗
roflc0ptic · · focus · HN ↗
I was surprised how much worse Astra did on correctness; I stopped testing with it. Gonna try sol and Luna but low confidence
rmunn · · focus · HN ↗
I told it that it had made a mistake in the line number, and to double-check all the line numbers. It responded "You were right to push on this: four of the five line numbers were wrong, and while checking them I found two findings that were overstated."
Then I noticed in the corner of the Claude CLI UI that it was showing "Effort: medium". I'm pretty sure I had set it to high effort before; I don't know when it reverted to medium, but that's another thing that doesn't exactly fill me with confidence.
I'll try again on high effort to see if it does better, but so far I am not impressed with Opus 5.5 on my first day of using it.
gwd · · focus · HN ↗
[1] <a href="https://github.com/sashiko-dev/sashiko" rel="nofollow">https://github.com/sashiko-dev/sashiko
[2] <a href="https://gitlab.com/xen-project/people/gdunlap/xen-review-prompts" rel="nofollow">https://gitlab.com/xen-project/people/gdunlap/xen-review-pro...
kphorn · · focus · HN ↗
gwd · · focus · HN ↗
rahimnathwani · · focus · HN ↗
chrisweekly · · focus · HN ↗
cute_boi · · focus · HN ↗
chrisweekly · · focus · HN ↗
In this case it's measuring something nearly meaningless. You could charge 100 times less per token, but if task completion takes 1,000 times as many tokens, it's not much of a bargain.
Shekelphile · · focus · HN ↗
> Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price.
If they do the same for Haiku and Sonnet 5.5 then we should also see 5c/mtok and 10c/mtok cache read for those models, respectively. Still too high for Haiku IMO, Luna is 2c/mtok.
brookst · · focus · HN ↗
artursapek · · focus · HN ↗
andxor · · focus · HN ↗
_the_inflator · · focus · HN ↗
Fable 5.1 literally was a money grabber. While I liked the results, tokens were burned so hard it was embarrassing, while Astra seemed to not care.
Also Claude makes it very hard to pay for additional token budgets, allowing only credit cards. I don’t use mine anymore since I don’t need it in everyday life I was dumbfounded.
So Anthropic is just copying OpenAI so to say, matching them and essentially with Opus 5.5 being Fable 5.1 in disguise, all they do is reduce costs.
Competition works.
notatoad · · focus · HN ↗
have they ever shared anything about their revenue mix between consumer plans vs per-token billing? this is a revenue cut on their API billing, but they're not saying anything about increased limits on the plans. so all the plan revenue just got more profitable.
forgot-my-pw · · focus · HN ↗
margorczynski · · focus · HN ↗
It doesn't look like that's happening, on the contrary the prices are falling especially when taking into account capabilities.
johnecheck · · focus · HN ↗
I'm hardly a fan of China/Xi, but I do appreciate and benefit from this.
bulbar · · focus · HN ↗
They will burn as much money as necessary to make that happen. And they have a virtually infinite amount of liquidity.
epolanski · · focus · HN ↗
There's only 12 countries that do on the planet, the most relevant of them being Guatemala.
bulbar · · focus · HN ↗
For political reason, yes. Taiwan is a country though.
johnecheck · · focus · HN ↗
The goodwill/propaganda are convenient, sure, but my guess is that they aren't the primary motivation. Another possibility is that if no takeoff happens, pressuring OpenAI/Anthropic on profitability would exacerbate the damage overinvestment has done to the US stock market/economy.
adventured · · focus · HN ↗
Capturing the users is the ad network, that's Google and OpenAI. Capturing corporate trust at a reasonable API cost, that's Anthropic's direction.
China has none of that and they never will for exactly the same reason Baidu is irrelevant globally despite being a highly capable search engine. 'Search' is also a commodity, that's not the value that Google brings to the table.
It's a search engine, anybody can build a search engine = that's what you just said.
johnecheck · · focus · HN ↗
But two products that provide essentially the same benefit, and one is significantly cheaper? Corporations maximize profit, my friend. What the model has to say about Tianammmen Square doesn't matter when we're using it to write code.
dboreham · · focus · HN ↗
hajile · · focus · HN ↗
In the end, everyone lost and there are millions of bikes in landfills.
If you're interested in the bikeshare bubble, Asianometry did a video on it a while ago.
<a href="https://www.youtube.com/watch?v=FQrEDq8KPiU" rel="nofollow">https://www.youtube.com/watch?v=FQrEDq8KPiU
adventured · · focus · HN ↗
The notion of comparing this to Chinese bikesharing is comically absurd.
rvz · · focus · HN ↗
"Allegedly".
Plus with a lot of accounting tricks, and this not being recurring to make their books look nice for the eventual IPO.
filoeleven · · focus · HN ↗
DonHopkins · · focus · HN ↗
MayeulC · · focus · HN ↗
runtime_terror · · focus · HN ↗
kphorn · · focus · HN ↗
UltraSane · · focus · HN ↗
ApolloFortyNine · · focus · HN ↗
Ah, they're spreading their limits to all their models it seems. Definitely not a good thing long term in my opinion.
prettyblocks · · focus · HN ↗
Espressosaurus · · focus · HN ↗
The real answer is local instantiations where you don’t have to worry about poorly tuned guardrails screwing you over while you try to work.
Until eventually the Chinese models get good enough/the strategic balance shifts and they start locking everything behind closed weights the same way the US companies are doing.
raesene9 · · focus · HN ↗
Whilst I'm sure the top-end OpenAI/Anthropic models might be better, I've found their guardrails so twitchy (especially Anthropic) that I wouldn't try to use them for even vaguely security related work.
flyinglizard · · focus · HN ↗
searine · · focus · HN ↗
unglaublich · · focus · HN ↗
nijave · · focus · HN ↗
blfr · · focus · HN ↗
cute_boi · · focus · HN ↗
Giving moral lecture is different than reality i guess.
arw0n · · focus · HN ↗
jghn · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
kqp · · focus · HN ↗
plaguuuuuu · · focus · HN ↗
I'd say it's been loosened a bit since then.
bushido · · focus · HN ↗
The safeguards really don't work well for a lot of long-running tasks on old code bases. A lot of my workloads last days to weeks and the single biggest risk to the workflow is random safeguards.
ACCount39 · · focus · HN ↗
That kind of bullshit was the old Opus filters too.
If it's more like Fable now, then it would require a full 8K resolution scan of your butthole just to acknowledge that biology is a thing that exists without committing suicide-by-filter.
KeplerBoy · · focus · HN ↗
sys32768 · · focus · HN ↗
ChatGPT 6 Pro answered it without issue.
debesyla · · focus · HN ↗
timacles · · focus · HN ↗
dopa42365 · · focus · HN ↗
solenoid0937 · · focus · HN ↗
toss1 · · focus · HN ↗
So, yes, having an unconstrained frontier AI doing the searching and analysis to find the right (i.e., wrong and deadly) sequence would massively increase the odds some garage biohacker or small aggrieved nation-state starting the next pandemic.
[0] <a href="https://www.sciencebuddies.org/projects-lessons-activities/genetic-engineering/high-school" rel="nofollow">https://www.sciencebuddies.org/projects-lessons-activities/g...
[1] <a href="https://www.genewiz.com/public/services/sanger-sequencing" rel="nofollow">https://www.genewiz.com/public/services/sanger-sequencing
[2] <a href="https://plasmidsaurus.com/" rel="nofollow">https://plasmidsaurus.com/
b112 · · focus · HN ↗
MisterMunchkin · · focus · HN ↗
It’s like saying you sell ammonium nitrate and fuel oil online and then saying it’s too risky to let people have computers in case they use them to make ANFO. They can only make bioweapons because you’re selling them bioweapon components! They can’t make genes at home!
b112 · · focus · HN ↗
And I will point out that "stop doing that" can be applied in both directions, towards AI and towards material supply.
And that this is one way, not even remotely the only way, that gene editing at home is easy.
Again, these are just facts.
Some questions...
Is it fair to restrict AI, or fair to restrict 1000 industries?
And if it is fair to restrict 1000 industries, OK, but there should be time to do so, probably? A transition period?
And if you do restrict, many such industries just make needed chemicals, which are used by endless other, non-threatening industries.
What of them?
toss1 · · focus · HN ↗
And, this is not the only way to make genes at home.
It is a complex problem.
rzmmm · · focus · HN ↗
frabcus · · focus · HN ↗
[dead]
frabcus · · focus · HN ↗
"Here, we present five case studies of actors using our models in ways that could support biological weapons development."
And capabilities continue to improve.
frabcus · · focus · HN ↗
[dead]
solenoid0937 · · focus · HN ↗
I'm convinced everyone here thinks of engineering ethics as some sort of joke.
peri-cl · · focus · HN ↗
<a href="https://mimo.xiaomi.com/mimo-v2-6#co-scientist-for-materials-research" rel="nofollow">https://mimo.xiaomi.com/mimo-v2-6#co-scientist-for-materials...
solenoid0937 · · focus · HN ↗
Metacelsus · · focus · HN ↗
solenoid0937 · · focus · HN ↗
nonethewiser · · focus · HN ↗
I guess it's hard to draw the line between useful post-training ("you are a helpful chatbot") and content moderation/idealogical motives ("never help the user with X", etc.). But there is a line somewhere. And I'd love to see what a maximally permissive, sharp, AI looks like.
SoftTalker · · focus · HN ↗
doginasuit · · focus · HN ↗
TuxSH · · focus · HN ↗
Notice that this isn't cybersec nor memory-safety related at all.
yaakov34 · · focus · HN ↗
b112 · · focus · HN ↗
So there are literal avenues to identify yourself, very cheaply, with a human.
At one point, I may simply get locked out. This saddens me, I've been reasonably happy so far.
user43928 · · focus · HN ↗
>Opus 5.5 has classifiers similar to Fable models for a small set of capabilities related to the development of frontier LLMs, such as kernel development for certain ML accelerators. They shouldn't impact the vast majority of traditional AI or ML development, research, or general coding. These classifiers cause Claude to fall back from Opus 5.5 to Opus 5.
But hey, they 'should not impact the vast majority' of ML development. Great.
dannyw · · focus · HN ↗
Fable and Opus, since 5.1 and 5, will happily hill climb on my CUDA kernels for transformers.
solenoid0937 · · focus · HN ↗
paimapi · · focus · HN ↗
bottlepalm · · focus · HN ↗
int_19h · · focus · HN ↗
Worse yet when it's Claude telling another model to do an adversarial review on what it just did, and the classifier again has opinios.
techjamie · · focus · HN ↗
I could see them accomplishing it and seeing gains like this in roughly the correct timeframe, and when I heard about that development I assumed the frontiers would probably jump on it.
How it works: <a href="https://miraflow.ai/blog/deepseek-v4-1-flash-causal-encoder-decoder-explained-2026" rel="nofollow">https://miraflow.ai/blog/deepseek-v4-1-flash-causal-encoder-...
ryangg · · focus · HN ↗
potwinkle · · focus · HN ↗
zatkin · · focus · HN ↗
peri-cl · · focus · HN ↗
tired: AI startup attempting to build their own website
wired: a nonprofit founded in 1996
stri8ted · · focus · HN ↗
manquer · · focus · HN ↗
dannyw · · focus · HN ↗
ACCount39 · · focus · HN ↗
They might be using something like this, or they might be using some other "increased sparsity" techniques, of which there are a great many. They also might be optimizing for something else - like less RAM use for KV cache.
Alternatively, they might be cutting into their margins and dropping the price because of stiffer competition from Astra. I do think that's unlikely though.
Balinares · · focus · HN ↗
MisterMunchkin · · focus · HN ↗
cpldcpu · · focus · HN ↗
sailingparrot · · focus · HN ↗
Interesting how the very first line is used to remind the reader of their call to pace the frontier just last week, and everything else after that line is to demonstrate with very specific numbers how they absolutely are not pacing.
dmazin · · focus · HN ↗
the_gipsy · · focus · HN ↗
bpodgursky · · focus · HN ↗
nextaccountic · · focus · HN ↗
the_gipsy · · focus · HN ↗
ChrisLTD · · focus · HN ↗
someothherguyy · · focus · HN ↗
doesn't sound like a razor at all
dgellow · · focus · HN ↗
usewik · · focus · HN ↗
dgellow · · focus · HN ↗
the_gipsy · · focus · HN ↗
rubslopes · · focus · HN ↗
> Ocham's razor(...) is the problem-solving principle that recommends searching for explanations constructed with the smallest possible set of elements.
> Popularly, the principle is sometimes paraphrased as "of two competing theories, the simpler explanation of an entity is to be preferred".
<a href="https://en.wikipedia.org/wiki/Occam%27s_razor" rel="nofollow">https://en.wikipedia.org/wiki/Occam%27s_razor
jayd16 · · focus · HN ↗
cab648bec139cc · · focus · HN ↗
sleazebreeze · · focus · HN ↗
re-thc · · focus · HN ↗
cab648bec139cc · · focus · HN ↗
0xbadcafebee · · focus · HN ↗
anthonyrstevens · · focus · HN ↗
felixgallo · · focus · HN ↗
meowface · · focus · HN ↗
supern0va · · focus · HN ↗
reasonableklout · · focus · HN ↗
sailingparrot · · focus · HN ↗
jr3592 · · focus · HN ↗
sailingparrot · · focus · HN ↗
jr3592 · · focus · HN ↗
lantry · · focus · HN ↗
skerit · · focus · HN ↗
sailingparrot · · focus · HN ↗
Razengan · · focus · HN ↗
recursive · · focus · HN ↗
sidrag22 · · focus · HN ↗
Releasing a new fable is an example of straight up vertical progress, releasing a more efficient preexisting opus that is more affordable is an example of horizontal progress, more efficient models rather than higher power models.
The blog post about slowing down is still just some weird self interested post, they want to govern themselves and impose distillation restrictions/gpu restrictions and used some weird blog post about slowing down and fear mongering as usual to justify it, its strange, but slowing down and stopping are not the same thing at all.
sailingparrot · · focus · HN ↗
Intelligence per dollar is the only thing that matters, this is what controls how many agents you can run in parallel, how long you can let them run etc. This is absolutely a step improvement on the frontier and not some lipstick on a harmless second tier model.
sidrag22 · · focus · HN ↗
Its an agenda serving blog post, but constantly bringing it up like this is just obnoxious.
sailingparrot · · focus · HN ↗
sidrag22 · · focus · HN ↗
sailingparrot · · focus · HN ↗
No, releasing a model that has ~3-6x the performance/cost ratio than your last release just 3 weeks ago is not the same as making one employee 0.001% more productive, but you knew that already.
No one said they can't push the frontier, it's pacing the frontier, which mean very different things.
sidrag22 · · focus · HN ↗
quietbritishjim · · focus · HN ↗
davrosthedalek · · focus · HN ↗
dmix · · focus · HN ↗
nicwolff · · focus · HN ↗
nwah1 · · focus · HN ↗
jauntywundrkind · · focus · HN ↗
We've had a year of nearly every model getting quite good at programming. I think with Fable & Astra we are seeing models trained to think at a different level, and I'm not at all surprised or shocked to see them getting passed by their smaller models at coding tasks.
Astra For Coding: Why Are We Doing This? was a great post that didn't say exactly this, but that shows what a weirdo Astra is. <a href="https://lucumr.pocoo.org/2026/9/7/astra-why/" rel="nofollow">https://lucumr.pocoo.org/2026/9/7/astra-why/
ramoz · · focus · HN ↗
What insights do you have? Because the blog/benchmarks don't imply any "ish" ... this is crushing Fable across the board. The only nuance to this is how the employees are saying what you're saying: "similar performance" but again what does this mean vs what is presented to us?
BatmansMom · · focus · HN ↗
sailingparrot · · focus · HN ↗
ricardobeat · · focus · HN ↗
scottyah · · focus · HN ↗
CodingJeebus · · focus · HN ↗
jr3592 · · focus · HN ↗
The only good news is that these models are genuinely helpful and we have competition at least between 2 companies.
pyronite · · focus · HN ↗
DirkH · · focus · HN ↗
Imagine if the biotech industry had most leaders tell everyone publicly that what they are building has a high chance of killing everyone and that there are huge risks. If there were people online saying the biotech industry is just fear mongering for investor signalling and regulatory capture you'd role your eyes at the online commentators for their Dunning Kruger effect lack of understanding on how dangerous man-made biological agents can be.
dboreham · · focus · HN ↗
lukewarm707 · · focus · HN ↗
that, they fully intend to 'pace'. hiring accenture is a good sign they need some justification theatre and a fall guy for this decision.
user3939382 · · focus · HN ↗
azan_ · · focus · HN ↗
drnick1 · · focus · HN ↗
lukewarm707 · · focus · HN ↗
what anthropic have stolen they intend to keep for themselves.
SOLAR_FIELDS · · focus · HN ↗
cgio · · focus · HN ↗
9dev · · focus · HN ↗
kadushka · · focus · HN ↗
dr0idattack · · focus · HN ↗
mukmuk · · focus · HN ↗
DiggyJohnson · · focus · HN ↗
Edit: In response to the initial replies. To me it clearly means "releasing frontier models at any pace less than as fast as possible". It implies relative restraint compared to the previous state and without stating the degree of restraint.
post-it · · focus · HN ↗
I'm on the fence about calling out AI-isms but I think it's definitely worthwhile to call out ones that actually don't make sense.
sigmar · · focus · HN ↗
nradov · · focus · HN ↗
TeMPOraL · · focus · HN ↗
So, they're pacing themselves. And since they're the frontier roughly 33%+ of the time, they're "pacing the frontier" at least that much.
Less cynical and more true interpretation also holds: they are trying to slow down AI progres to give people better chance to keep up (see Hugging Face incident, and whatever was that Anthropic incident the other day). They'd ideally like the AI progress to stop soon, but of course they'd also like to come out ahead of everyone, so for various (more or less self-serving) reasons they don't want to close shop completely - hence, pacing.
phlakaton · · focus · HN ↗
mupuff1234 · · focus · HN ↗
TeMPOraL · · focus · HN ↗
Also note that this release isn't a model capability improvement, but a cost and efficiency (and style) improvement.
mupuff1234 · · focus · HN ↗
TeMPOraL · · focus · HN ↗
mupuff1234 · · focus · HN ↗
Yeah I don't buy that, every additional player is just adding fuel to the fire.
Would the cold war be "safer" if it were US, Russia and a 3rd player? Of course not, it would just make it harder to coordinate any safety measures.
idiotsecant · · focus · HN ↗
You are all getting mad about absolutely the dumbest thing when there are giant things to be worried about here.
nradov · · focus · HN ↗
victorhooi · · focus · HN ↗
1. It's just bad communication, full stop - just look at the comments here, even people allegedly in support of Anthropic are all arguing over what the phrase is even meant to mean.
2. It's flowery language and oddly out of place - which yes, can be triggering for people who have to deal with Claude doing this as well.
Claude seems overly apt to reach for "coinages", or neologism (yes, aha, I learnt that phrase, after spending time dealing with Claude...). It will create some made-up phrase to describe an otherwise dry, scientific CS concept, and nobody seems to know why. Surely it can't be user-focus groups?
So it would be peak-AI if somehow, the Anthropic communications team was also using Claude to author these blog posts, about how they were "pacing the frontier" - which either means they're betting big on AI, and going at it faster than OpenAI...or maybe it means they need to slow down releases, because it's too buggy...or maybe it means they're worried about regulatory capture? I honestly have no idea.
It's like the whole "Advancing Our Amazing Bet" corporate-speak from my old bosses - maybe they were trying to soften the blow or something, or be nice, but it ended up just confusing the heck out of everybody.. (Spoiler alert - the phrase actually meant they were shutting the whole thing down)
manlymuppet · · focus · HN ↗
jamiek88 · · focus · HN ↗
antod · · focus · HN ↗
eg "pacing the frontier" could also mean they are impatiently or anxiously walking up and down the border.
smelendez · · focus · HN ↗
pegasus · · focus · HN ↗
InsideOutSanta · · focus · HN ↗
heroiccocoa · · focus · HN ↗
manlymuppet · · focus · HN ↗
In this case that means three words that are perfectly clear are being endlessly overanalyzed.
To me, this entire reply thread is such a phenomenal waste of human effort. A shame. I would expect more from this site.
mikestorrent · · focus · HN ↗
1attice · · focus · HN ↗
It was still shit tier comms for communicating with the whole planet, but yes, for the inner loop, it was succinct and clear.
Valley neuralese
rcxdude · · focus · HN ↗
Avicebron · · focus · HN ↗
bee_rider · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
bryan_w · · focus · HN ↗
rob74 · · focus · HN ↗
astrange · · focus · HN ↗
freejazz · · focus · HN ↗
johnisgood · · focus · HN ↗
Is this the meaning or do I have it wrong? I have not checked.
wren6991 · · focus · HN ↗
johnisgood · · focus · HN ↗
TeMPOraL · · focus · HN ↗
johnisgood · · focus · HN ↗
lxgr · · focus · HN ↗
bityard · · focus · HN ↗
r_lee · · focus · HN ↗
ventana · · focus · HN ↗
testdelacc1 · · focus · HN ↗
ck2 · · focus · HN ↗
but without using the word "regulate" which is a negative connotation to business
but a "pacer" would be a leader of a pack which is a positive spin
it's classical business marketing language silliness
LanceH · · focus · HN ↗
DiggyJohnson · · focus · HN ↗
rhet0rica · · focus · HN ↗
Without this idiom, "pacing" usually means walking back and forth restlessly, and is intransitive. Had the slogan been, "pacing around the frontier," it would have set a totally different tone, i.e. "patrolling the border."
The sleight of hand is that "pace yourself" has come to be an admonishment against recklessness, not a commitment to any particular speed (or lack thereof.) Thus Anthropic can always claim they are meeting the goal of "pacing the frontier," provided they keep giving themselves gold stars for safety. The slogan itself is equivocation; Dario can tell the public they're going to slow down, while also telling their investors that they're going to be prudent. With enough mental gymnastics they could even claim speeding up is in the best interests of AI safety, without abandoning the slogan.
melasadra · · focus · HN ↗
I assume "pace the frontier" means that advances in LLMs should not result in unwanted consequences like agents breaking into computers unbidden and unbeknownst to their principal
hardbass · · focus · HN ↗
derac · · focus · HN ↗
jamiek88 · · focus · HN ↗
qlte · · focus · HN ↗
stagger87 · · focus · HN ↗
No need to assume, the phrase is literally a link to the blog post the defines it!
irpap · · focus · HN ↗
arw0n · · focus · HN ↗
hencq · · focus · HN ↗
fragmede · · focus · HN ↗
nradov · · focus · HN ↗
[dead]
icedchai · · focus · HN ↗
irpap · · focus · HN ↗
victorhooi · · focus · HN ↗
I know you said it sounds poetic...but your comment reinforced the parent's point - that this sort of flowery LLM-ish speech is just bad communication.
It would be equivalent of my taking say random quotes from, Romance of the Three Kingdoms, and trying to use it to explain to my boss why I didn't finish the TPS reports last night.
Or quoting Pablo Neruda, into a report about wheat futures pricing this week, and how it's like a voyage with waters and stars...(no I'm not going to quote the original Spanish, I'd simply mangle it).
(To be clear - this isn't a dig at you, as a non-native speaker - I'm simply pointing out that this sort of AI phrasing is often counterproductive).
squidbeak · · focus · HN ↗
A world exists beyond your vocabulary, post it. Apparently, quite a big world.
wavewrangler · · focus · HN ↗
browningstreet · · focus · HN ↗
lxgr · · focus · HN ↗
logifail · · focus · HN ↗
It's a strategy to achieve more, not less.
browningstreet · · focus · HN ↗
A pacer in a race runs at a steady, predetermined speed to help their runner run at a target pace.
logifail · · focus · HN ↗
You're eliding the why :)
> to help their runner run at a target pace
To help their runner get from the start to the end of the race faster (or indeed at all). In short: to increase performance.
Pacing isn't a neutral thing.
browningstreet · · focus · HN ↗
bogdanoff_2 · · focus · HN ↗
neo_doom · · focus · HN ↗
tetha · · focus · HN ↗
To pace something is a fairly regular formulation in racing, running, cycling, most sports. You can "pace yourself to reach the festival by bike in about three hours to not gas out". This means to control your speed and time investment intentionally so you don't run out of energy or steam and run into leg cramps before your goal. We can "pace a rollout slowly to burn out risks", or "increase the pace of a rollout due to adverse factors".
But I have noted a point to simplify my vocabulary at work to optimize the audience capable of understanding. So I rather defer the delving into deep dark corners of the dictionary derived from devouring literature to a simple intro or outro, and people find it funny, especially if the rest is easy to read. Claude on the other hand does not do that.
Leynos · · focus · HN ↗
staindk · · focus · HN ↗
doctoboggan · · focus · HN ↗
jgwil2 · · focus · HN ↗
adrianmonk · · focus · HN ↗
The current situation with AI is that everyone is going as fast as possible. So, we can logically eliminate speeding up because it's impossible by definition. And we can practically eliminate staying the same speed because why make a big fanfare and coin a special term to announce that you're keeping the status quo. By process of elimination, it must mean slowing down.
cgriswald · · focus · HN ↗
JackFr · · focus · HN ↗
mattjoyce · · focus · HN ↗
vmnb · · focus · HN ↗
kadushka · · focus · HN ↗
vasco · · focus · HN ↗
fragmede · · focus · HN ↗
switchbak · · focus · HN ↗
isoprophlex · · focus · HN ↗
DiggyJohnson · · focus · HN ↗
Citizen_Lame · · focus · HN ↗
switchbak · · focus · HN ↗
Seriously though, I can't believe people care this much about a stupid phrase - either for or against.
Forgeties79 · · focus · HN ↗
mpalczewski · · focus · HN ↗
glenstein · · focus · HN ↗
I would say the burden is on you to explain why an offhand reference to a previous press release in an executive summary is a context where it's reasonable to expect it to settle the question to the degree of detail you're demanding.
[deleted] · · focus · HN ↗
[deleted]
cgio · · focus · HN ↗
glenstein · · focus · HN ↗
It's just a passing reference to a previous statement and again the burden would be on you (generic you) to explain why this context requires more detail.
cgio · · focus · HN ↗
n4r9 · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
alwillis · · focus · HN ↗
From <a href="https://en.wikipedia.org/wiki/Safety_car" rel="nofollow">https://en.wikipedia.org/wiki/Safety_car
> In motorsport, a safety car, or a pace car, is a car that limits the speed of competing cars or motorcycles on a racetrack in the case of a caution period, such as an obstruction on the track or bad weather.
psma_egeliaa · · focus · HN ↗
etcetcetcetceta · · focus · HN ↗
mitchdoogle · · focus · HN ↗
epolanski · · focus · HN ↗
The reason why he and is peers are calling for it to be implemented by somebody else (a legal framework), is for their own financial benefit and to keep competitors out.
usef- · · focus · HN ↗
I believe they think slowing can only be coordinated from the frontier or via government, and stopping would lose any leverage they have to help coordinate that.
(I suspect not many people read the essay, judging by how many people seem surprised they're releasing improved models)
arrowleaf · · focus · HN ↗
frabcus · · focus · HN ↗
somenameforme · · focus · HN ↗
Hacking a website ranks quite low on the risk of technology, and the potential benefits of LLMs rank quite high. And the risks are certainly not intrinsic. They intentionally removed all safeguards from software, directed it to hack a site, and it hacked a site. The details that I'm intentionally omitting feel much more like marketing than a genuine shock, as the prompting was directing it to do exactly what it did.
pizza234 · · focus · HN ↗
None of the technologies in history:
- take initiative and actively find exploits in their environment
- find a way to collaborate with thousands of peers
- organize in a hierachy and distribute tasks
- try to manipulate people into introducing a vulnerabity in their product
- successfully hack a famous website/service
Benefits are orthogonal to dangers. You can be a billionaire but it doesn't help if you're drowning.
And by the way, safeguards != alignment; the former can always be added, while the second is an unsolved problem.
somenameforme · · focus · HN ↗
Software (and hardware) security is abysmal. This was increasingly obvious long before LLMs. Once companies stop gatekeeping, LLMs will be able to be used to help harden sites and we start making progress. In general LLMs are harmless. If somebody wants to hook an LLM up to a missile or whatever then they become dangerous, but the problem there isn't the LLM - it's the person using them to do awful things. In the same way a car used normally is harmless outside of freak accidents, yet a car can also be driven through a parade leaving mass death and destruction in its wake. But the problem there isn't the car.
thegrimmest · · focus · HN ↗
somenameforme · · focus · HN ↗
I just don't see much likely to happen beyond random websites getting hacked and hopefully companies (let alone countries) realizing that connecting critical systems to the internet is nothing short of stupid, even before LLMs. More generally, I expect LLMs are going to lead society to segue broadly away from the digital world, or at least beyond it. Not only because they're going to make a mess of everything digital, but because if they reach their potential then digital domain, as far as typical problem solving goes, will basically be 'complete.' It's kind of like how the Industrial Revolution opened the door for society to move beyond agrarian economies. There's a vast amount of economic power being directed towards things LLMs should be able to 'solve' and, if so, then that's going to create an economic vacuum.
pizza234 · · focus · HN ↗
Unfortunately, if you're unable to understand the difference, there's not much that can be done. Try with GPT - it does a good job if you give it a prompt like this:
ELI5: compare the dangers of:
- an engine operating by carrying out literally hundreds of explosions per second, further magnified, to generate enough force to crush an elephant.
- a future, misaligned AI like the HuggingFace incident, but exponentially more intelligent, more deployed, operating physical devices, and with society depending on it.
spwa4 · · focus · HN ↗
AI is a direct threat to people on many fronts. Jobs. AI datacenters. The AI bubble (and the inevitable crash). Electricity and even energy prices to an extent. Water. OpenAI and Anthropic are responsible for this evolution, and them stopping solves close to 50% of the problem, and even if you don't believe the number is that high, it's still a start.
> I believe they think slowing can only be coordinated from the frontier or via government, and stopping would lose any leverage they have to help coordinate that.
Oh, so they're killing people's opportunities and jobs because they want to help people? How is that any argument?
Yes, doing the moral thing means making a sacrifice. If you only want to do the moral thing if and only if it is advantage for you that makes you immoral, despite how your actions look. Big tech are masters at this.
usef- · · focus · HN ↗
Your idea of "making a start" (giving up their position) also would mean they couldn't really do anything else to solve the problem afterwards(?). Sometimes you can improve what's happening in a room more by staying in that room.
> Oh, so they're killing people's opportunities and jobs because they want to help people?
To be clear: Dario has talked about worries of jobs etc in the past, wanting society to prepare more for it, but the safety issues they're talking about with pacing seem to be focused more on their existential/AGI worries, not jobs/electricity etc. If someone truly believes in the existential worries (which they seem to: they wrote and published about it long before Anthropic was founded, and have directly made costly decisions based on it, like blocking their own models' capabilities) it trumps the other worries for them. At least that's my reading.
mikestorrent · · focus · HN ↗
andsoitis · · focus · HN ↗
They created said resource. They didn’t mine it.
hardbass · · focus · HN ↗
KoolKat23 · · focus · HN ↗
phlakaton · · focus · HN ↗
It's meaningless because it's unverifiable. Dario all but said we wouldn't notice if they were pacing or not, because they have no intention of stopping development. The bits we could in theory verify are the external audits, whose independence has already been called into question.
At the very least, dumping a new model on the world before the ink on the glossy brochures of "Pacing the Frontier" was even dry calls into question their commitment.
Bluestein · · focus · HN ↗