Grok 4.7
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
Grok 4.7
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
ls1911 · · focus · HN ↗
simianwords · · focus · HN ↗
The personality is bland and it doesn’t work nearly as hard or even tries to help.
artemonster · · focus · HN ↗
xutopia · · focus · HN ↗
artemonster · · focus · HN ↗
Paracompact · · focus · HN ↗
xutopia · · focus · HN ↗
There are ample reasons to believe that Elon Musk is running his mother's account and that the photos weren't even real putting in question that he even had a birthday party.
If you ask Grok about what this means it will always take Elon's defence. It will vehemently deny that Elon would be capable or willing participant of such a thing even if you point out that he faked being a world class gamer, buying accounts that had done all the work and showing none of the skills when live-streaming.
jpadkins · · focus · HN ↗
Yeah, I'm sure the guy has time to run his mom's social media account. What's your reasons or evidence? Did you consider maybe it's a social media manager one of them hired running the account?
slowin · · focus · HN ↗
Capricorn2481 · · focus · HN ↗
I don't use Grok, but do you want your LLM to have a personality? "Personality" is exactly what people don't like about Claude.
Razengan · · focus · HN ↗
nython · · focus · HN ↗
Razengan · · focus · HN ↗
longdong1 · · focus · HN ↗
Razengan · · focus · HN ↗
created: 9 minutes ago
That could have been said just as perfectly well from a main account my guy/guyette
Razengan · · focus · HN ↗
ethagnawl · · focus · HN ↗
Until you ask it to start generating horrific imagery and then it's best in class.
Shiggy_ · · focus · HN ↗
The value of the internet is that people can share whatever they want, and use software how they want. This will mean that some people will abuse that. This is the tradeoff of a free society.
raincole · · focus · HN ↗
Sounds like a plus. Guess I will give Grok another try...
moojacob · · focus · HN ↗
Given that the decrease in their margin and the fact they delayed the release of Grok 4.7 almost two weeks past the original date, XAI must not have been happy with the results for 4.7. And XAI also waited the day before Opus 5.5 is rumored to launch. I imagine Opus 5.5 will blow Grok 4.7 out of the water benchmark wise.
However, I have become skeptical of benchmarks. Grok 4.5 solved some issues setting up a buildroot system that Fable 5 couldn't do. I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby.
My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish. Or even GPT, which doesn't have Claude's ticks but definitely likes to handwave explaining technical concepts. Still, nothing beats Claude 3.5 and 4 with explaining since it seems all models have regressed. I wonder if Grok 4.7 will also regress with English because of all the RL.
jasonjmcghee · · focus · HN ↗
That being said, I currently prefer Sol / Astra to Opus / Fable as I find both to be a better cost payoff to me.
vessenes · · focus · HN ↗
vintermann · · focus · HN ↗
svachalek · · focus · HN ↗
vintermann · · focus · HN ↗
boc · · focus · HN ↗
hombre_fatal · · focus · HN ↗
Astra blows up limits way too fast to be useful for anything beyond really hard bugs when I'm already running out of Claude usage. But Sol (low) is my other bread and butter workhorse agent next to Opus.
jitl · · focus · HN ↗
boc · · focus · HN ↗
jitl · · focus · HN ↗
dumberquestions · · focus · HN ↗
moojacob · · focus · HN ↗
And plus, token efficiency isn't enough to tell you much if we are going to play that game. You might as well just measure how long it takes to complete a task. GPT is super token efficient but takes longer than Grok sometimes because Grok is unpopular and they have more spare capacity to serve your request because Musk impulsively splurged on a giant datacenter.
user43928 · · focus · HN ↗
How representative that is of real world usage, I don't know.
In their benchmark GPT 5.6 Sol performs suspiciously poorly compared to the former models.
formerly_proven · · focus · HN ↗
The strength of the grokish dialect is similar to GPT, but weaker than claudish, but grokish is just closer to normal language on average.
Lucasoato · · focus · HN ↗
I totally agree, it’s like that as models become more intelligent, they are less understandable by most of people... but aren’t we humans doing the same?
Aperocky · · focus · HN ↗
When claude speak in convoluted mess, they are often going off on tangents in real work that you asked it to do, too.
fragmede · · focus · HN ↗
Just because something is difficult to understand doesn't mean it's fraud, although if someone is trying to dazzle you with clever words and names of institutions you recognize because they are selling you something, there's a good chance they're lying to you in order to get some money from you.
Pannoniae · · focus · HN ↗
Sure the explanation will oversimplify a lot but then you can expand it recursively if needed, you gotta start somewhere.
includenotfound · · focus · HN ↗
You just simplified most of the problems people work on down to cancer complexity. Ironic, isn't it?
That's also simply not the case, most people are building CRUD apps with some frontend code and some accessory stuff like build systems etc., which while complex, can still be expressed in very plain, easy to understand language for anyone who's a bit technical.
Does not excuse the Claude slop.
fragmede · · focus · HN ↗
Solving the problem right in front of you is easy. Stepping back and asking: is that a problem to be solved, is infinitely harder.
I did not use Claude to write my comment, so I don't know where that is coming from.
includenotfound · · focus · HN ↗
Aperocky · · focus · HN ↗
If the presenter can't divide a problem until it reaches a series of independently simple concepts, then there's usually something fishy going on.
menaerus · · focus · HN ↗
Aperocky · · focus · HN ↗
grababner · · focus · HN ↗
michaelmrose · · focus · HN ↗
[dead]
samuelknight · · focus · HN ↗
TomGarden · · focus · HN ↗
The weird thing is, that's not what AI models seem to be doing. The prose is just weird.
unshavedyak · · focus · HN ↗
This happens most though when the speaker doesn't (or care to) understand their audience.
Eg i find effective communication requires expertise in both the subject matter domain but also the reference of the listener. Eg in ELI5 framing, if you don't know what information 5yr olds are expected to know you'll do a poor job at an ELI5.
It often feels like Claude does poorly at both framing the response relative to what it "thinks" the listener knows, but also the prose is... sideways, just weird as you said.
pixl97 · · focus · HN ↗
cyanydeez · · focus · HN ↗
It is unsurprising that a LLM fails, without coaching, to effectively communicate.
TomGarden · · focus · HN ↗
If I don't grok an elaborate explanation, I can ask for clarification. If it's explained to me in an overly simplistic or unnuanced way, I'll walk away with a false sense of understanding.
That said, I'm sure we all have very different concentrations of these types of people and problems around us. I've definitely met some engineers who seem to actively try to make their language incomprehensible
MisterMunchkin · · focus · HN ↗
fearmerchant · · focus · HN ↗
Agreed. Do you think it's due to that EU issue of making AI text be identifiable?
flipthefrog · · focus · HN ↗
superjan · · focus · HN ↗
svachalek · · focus · HN ↗
As for the wording of the prompt, you're pretty on point, I created a custom output style targeting mostly the first two you have there. Some people have wording that demands a certain technical standard or uses fancy words to describe what to avoid, but I haven't seen evidence those work better than asking plainly and I suspect the opposite: LLMs mimic the user to a degree so talking to it in terms of technical specifications and fancy words is an invitation to get them back.
thesmtsolver2 · · focus · HN ↗
Part of intelligence is knowing your audience and communicating efficiently.
cruffle_duffle · · focus · HN ↗
Bingo! And on this axis many SOTA models fail miserably. These things are acting on my behalf under my direction. All the supposed intelligence in the world means fuck-all if nobody can understand it.
And like somebody else said… when meat-based humans talk like Claude does, it almost always means they either don’t understand what they are talking about, or are actively trying to conceal something and are a fraud. Not always, but almost always.
yread · · focus · HN ↗
r_lee · · focus · HN ↗
this is specifically an Anthropic problem, maybe due to their heavy use of Claude to train Claude itself?
atniomn · · focus · HN ↗
sscaryterry · · focus · HN ↗
7734128 · · focus · HN ↗
fatata123 · · focus · HN ↗
sscaryterry · · focus · HN ↗
moojacob · · focus · HN ↗
Fable 5.1 is not there quite there yet.
They need to get that Sonnet 3.5 magic back.
rfgplk · · focus · HN ↗
// `HANDLE` is an opaque kernel handle (kernel32 validates and returns 0/FALSE // on a non-console handle); every out-param is `&mut T` to a `#[repr(C)]` POD, // ABI-identical to the Win32 `LP` pointer (thin non-null). The reference type // encodes the only pointer-validity precondition, so `safe fn` discharges the // link-time proof. (`bun_windows_sys::kernel32` declares these with `mut`; // redeclared locally so the legacy-conhost cursor path below is plain calls.)
or
// Progress's terminal handle is the canonical `output::File` (vtable-backed // stderr/File from `OutputSinkVTable`). The duplicate `ProgressTerminalVTable` // from B-0 round 1 is removed; tty/ansi/winsize route through the new // `OutputSinkVTable` slots so `bun_core` stays T0 (no `bun_sys` dep).
from src/bun_core/Progress.rs
mrieck · · focus · HN ↗
r_lee · · focus · HN ↗
just reading this gives me a headache
imron · · focus · HN ↗
The Claudish is dead. Long live the Claudish.
WarmWash · · focus · HN ↗
AustinDev · · focus · HN ↗
jtwaleson · · focus · HN ↗
esafak · · focus · HN ↗
svachalek · · focus · HN ↗
moojacob · · focus · HN ↗
I am a huge fan of Gemini Pro for chat... gemini somehow just knows the most obscure stuff. I'll double check something Gemini said and find the source is deep inside a hard to access scientific paper. Google just has the best index of the internet.
WarmWash · · focus · HN ↗
It's best for brain storming, rabbit holes, and image recognition.
Let the big models do the heavy lifting for now.
ipsod · · focus · HN ↗
Even if you aren't coding, you really need to double check its answers. Flash 3.8 hallucinated a Keyence camera's max operating temperature for me, last week, and backed it up with "references".
It's still my favorite model for most non-coding stuff, though.
haellsigh · · focus · HN ↗
StilesCrisis · · focus · HN ↗
r_lee · · focus · HN ↗
xmorse · · focus · HN ↗
smashers1114 · · focus · HN ↗
_boffin_ · · focus · HN ↗
LPisGood · · focus · HN ↗
kekebo · · focus · HN ↗
MisterMunchkin · · focus · HN ↗
ffsm8 · · focus · HN ↗
Literally every one, even 1-2 prompts later it starts to go back
bel8 · · focus · HN ↗
tempest_ · · focus · HN ↗
It burns more tokens but is the only way to get tolerable text.
hungryhobbit · · focus · HN ↗
tempest_ · · focus · HN ↗
<a href="https://code.claude.com/docs/en/hooks-guide#agent-based-hooks" rel="nofollow">https://code.claude.com/docs/en/hooks-guide#agent-based-hook...
hungryhobbit · · focus · HN ↗
snapplebobapple · · focus · HN ↗
a2dam · · focus · HN ↗
First, after a while it's just as grating as Claudeish. Second, my hunch is that it constricts the actual thinking of the LLM, like the same way that Newspeak does in 1984. It shrinks the range of thought that can be expressed if used as an input.
I think the real way to do it is to have another Claude entirely deal with the user as a liaison, but to keep the thinking in whatever format it came in.
Latent space reasoning, if you think about it, is exactly this to a crazy degree: why even formulate a thought as words if you can just keep it as matmuls until the user needs it? And then, if the user needs it, have it always specifically formulated for the user by another LLM rather than constrict its range of thought? Anyway, that's my take.
StilesCrisis · · focus · HN ↗
TuxMark5 · · focus · HN ↗
nomel · · focus · HN ↗
I don't think this is true.
They have to express themselves as tokens. The meaning of those tokens doesn't have to be text. See any model that can handle images/video. Also, I don't think math, svg, etc, are "natural" language.
And, only the final expression is tokens. The intermediate layers, with the encoded concepts, aren't "natural language".
But, to address your concern, potentially: <a href="https://news.ycombinator.com/item?id=49758615">https://news.ycombinator.com/item?id=49758615
_puk · · focus · HN ↗
The model isn't limited to concepts that can be expressed in natural language.
It's only once the AI gets to the output layers that natural language comes back into play.
After all, they're all made out of weights[0].
0: <a href="https://maxleiter.com/blog/weights" rel="nofollow">https://maxleiter.com/blog/weights
lelanthran · · focus · HN ↗
How do we know for sure? We don't even know how the emergent properties we see actually emerged?
For humans we know for sure that people sometimes have concepts that they have no word for (the reason the phrase "It's on the tip of my tongue" is a phrase, after all).
We don't know this for LLMs. When it makes new phrases, it's always a mixup of two existing words hyphenated (aside, that also seems to be the limits of SOTA models creativity - join two unrelated words together with a hyphen).
LLMs never respond with "It's on the tip of my tongue" type responses, indicating it has a concept but cannot remember (or does not have) a word for that concept. Every human, pre-speech-age, has managed to express or convey concepts that they had no word for.
So, no. I'd need a citation, preferably multiple, that did the trials and found that a model can generate concepts for which it does not have any words for.
timacles · · focus · HN ↗
Dylan16807 · · focus · HN ↗
nomel · · focus · HN ↗
helloplanets · · focus · HN ↗
booty · · focus · HN ↗
1. Is natural language holding LLMs back by some %? 2. Is natural language serving as a hard gate that will prevent LLM intelligent progressing past some specific point?
The answer to 1 seems like an obvious yes to me.
Your thesis says the answer to 2 is "yes." That doesn't feel right to me. Think about all of the humans who have pushed various fields forward: Einstein, Newtown, Bach, whoever. If natural language doesn't prevent an entity from surpassing humans in one intellectual field, why would it prevent an entity from surpassing humans in all intellectual fields?
(To be clear, I'm not claiming superintelligence will or won't be achieved; I'm considering your specific thesis about whether or not natural language will be a hard gate)
dist-epoch · · focus · HN ↗
faeyanpiraat · · focus · HN ↗
indigo945 · · focus · HN ↗
By the way, how good is Claude's Hopi?
smashers1114 · · focus · HN ↗
I do think an infrastructure where another Claude retranslates the output would be better. Oftentimes I forget to put it in the actual prompt and when I receive back 8 paragraphs of Claudeish I ask for it then.
I would have to disagree that it gets as grating as Claudeish though. Its just direct and professional instead of ring-around-the-rosy clickbait.
powvans · · focus · HN ↗
“I would have to disagree that it gets as grating as Claudeish though.”
It’s hard to imagine anything more grating than Claudeish. To quote Rainer Wolfcastle, "My eyes! The goggles do nothing!"
oxidant · · focus · HN ↗
cjonas · · focus · HN ↗
mrandish · · focus · HN ↗
I've found that prompting any constraint on output (length, style, vocab, even simple formatting) not only places additional cognitive load on the model, which burns some of whatever cognitive budget is available, it will also often skew the output in other subtle and completely unrelated ways.
Since I found this artifact interesting, I did some pretty extensive experiments a couple months ago. The increased load is real, although it may not be apparent if you're not near any cognitive boundaries. The subtle skew, however, seems nearly ever-present regardless of load.
evulhotdog · · focus · HN ↗
mrandish · · focus · HN ↗
While extensive, my tests were just following my curiousity, not controlled, exhaustive or well-documented. I identified about a dozen prior sessions of varying length and complexity to test and downloaded them with a browser add-on. I then removed all other user prompt instructions except for the formatting instruction. A test would typically involve changing the wording of the formatting instruction ranging from brutally simple to detailed and complete, then starting a new session, seeding one of the test sessions and continuing it. To get a feel for baseline inter-session variation, I also tried running the exact same prompt/session multiple times back-to-back, at different times and on different days of the week.
Once I identified a promising prompt candidate, I'd make it the formatting instruction in my regular, daily-use prompt for a few days. I quickly got a feel for how seemingly minor user prompt variations impact response quality, compliance and tone across fresh sessions as well as those in various states of context rot, drift, decay and cliff (<--my ni cknames for the distinct flavors of session degradation than technical terms of art).
My overall conclusion was that every instruction, no matter how minor or unrelated it seems, has some, real impact on the model's cog load, attentional focus and/or attentional weight budget. Both how these impacts manifest and what causes more or less impact is often extremely counteriintuitive. To more fully understand this, I eventually, got to the point of testing null case variants, such as the entire user prompt being one sentence completely unrelated to text formatting or the session topic, like: "Don't reference the cartoon character SnagglePuss" (in a deep dive on ancient Sumerian clay tokens). Similarly, a simple one sentence prompt requesting something the model already always does naturally also has a cost (eg "Capitalize proper nouns"). As others have observed, heavy emphasis, absolute prohibitions or emotional weight in prompts also tend to have outsized impact in both skew (impacting unrelated output tone/style) and in accelerating session degradation. "Avoid referencing SnagglePuss when you can" would have equal compliance but fewer downside impacts than "NEVER reference the cartoon character SnagglePuss" in sessions starting to degrade.
There were also surprises, such as when I was scanning transcripts of an older, longer session and noticed the LLM was doing number formatting almost perfectly. On looking at the active user prompt at the time (I keep a log of every user prompt change I make for every model), it didn't even reference formatting at all. More experimentation showed it a result of the LLM gradually mirroring my consistent use of formatting structure in my prompts over a long session (in which I never mentioned anything about formatting). Unfortunately, that mirrored trait doesn't persist to new sessions and reaching that point requires a substantial number of rounds burning quite a bit of context window.
el_benhameen · · focus · HN ↗
SoMomentary · · focus · HN ↗
BatteryMountain · · focus · HN ↗
BatteryMountain · · focus · HN ↗
fr2029 · · focus · HN ↗
[dead]
guluarte · · focus · HN ↗
junon · · focus · HN ↗
neomantra · · focus · HN ↗
It’s been really productive and I’ve been asking my agents to communicate using it more and more. I believe it’s relieved my cognitive load a bit while working with them.
<a href="https://github.com/tinygo-org/tinygo/blob/dev/AGENTS.md" rel="nofollow">https://github.com/tinygo-org/tinygo/blob/dev/AGENTS.md
faangguyindia · · focus · HN ↗
Waterluvian · · focus · HN ↗
tk90 · · focus · HN ↗
Wonder if we'd benefit from a much more specialized + task-specific benchmarks to paint a clearer picture like this. A benchmark solely for frontend, ruby, hardware, etc.
dmix · · focus · HN ↗
algoth1 · · focus · HN ↗
shawabawa3 · · focus · HN ↗
fakwandi_priv · · focus · HN ↗
Googling it returns no matches but I think it was supposed to be “live viewer”?
r_lee · · focus · HN ↗
rayiner · · focus · HN ↗
johnsimer · · focus · HN ↗
pietz · · focus · HN ↗
petesergeant · · focus · HN ↗
Forgeties79 · · focus · HN ↗
I don't want to waste money because my calculator is cracking jokes. They don't deserve their paltry 5% marketshare or whatever it is they have currently. I'm not even getting into Musk as a person or the horrid things we've seen Grok spit out on twitter. I just don't trust his companies with my data and I have seen very little evidence that it's ever the best tool for the job. I'm sure those cases exist but I can't imagine it's worth it.
StilesCrisis · · focus · HN ↗
attentive · · focus · HN ↗
And like that grok4.7 cache reads are more expensive than sol's (at $0.40/mil).
quater321 · · focus · HN ↗
[dead]
giancarlostoro · · focus · HN ↗
I do wonder why a frontier model does this to be honest. It still does good coding wise, but it seems strange to me. r/Claude is full of "load bearing" jokes in every thread.
quater321 · · focus · HN ↗
[dead]
qaq · · focus · HN ↗
imron · · focus · HN ↗
Grok has its own feel too. It's not as bad as Claude, but one of the things that bugs me is that it is far too terse.
It regularly seems to come up with terms and descriptions for things in its chain of reasoning and then uses these terms in its output assuming you understand what it's talking about.
I find I often have to ask it to re-explain what it means.
iamflimflam1 · · focus · HN ↗
bilbo-b-baggins · · focus · HN ↗
I’m pretty sure the big bois don’t do it because it would undermine “confidence”.
Seeing a model output “Oh I should just delete blah. Wait blah is a production service, I shouldn’t touch that. Maybe I can gain access to blah? Oh the aws cli isn’t signed in to blah. I see kubectl has access to blah though! Wait, I should ask user permission first.”
Yeaaaaah. Thinking tokens are fuckin’ wild.
tyre · · focus · HN ↗
glub · · focus · HN ↗
Could be something very stupid like - "I don't have ffmpeg available here. Should I install it? No, I can't. I'll proceed doing something that will take me 100x more tokens and wall clock just to avoid adding a dependency." I can then just stop and say - you've got nix flake there, just add it.
That's impossible with western models. The only way is to ask why it did something stupid when it already spent your 50% of your weekly quota.
tyre · · focus · HN ↗
geetee · · focus · HN ↗
demibabs · · focus · HN ↗
greenavocado · · focus · HN ↗
imron · · focus · HN ↗
tempay · · focus · HN ↗
Zambyte · · focus · HN ↗
imron · · focus · HN ↗
octoberfranklin · · focus · HN ↗
I suspect that they specifically train Grok to be able to work well with military personnel -- speaking the way they speak: brief, to the point, efficient communication. Personally I really like this. Claude sounds like some demented clown from the marketing department.
imron · · focus · HN ↗
shuwix · · focus · HN ↗
[dead]
caminante · · focus · HN ↗
Anecdotally, I have noticed the same in the past week. It might just be anecdotal or driven by a long context window.
taspeotis · · focus · HN ↗
Separately have been using Grok 4.6 for a bit and it's also pretty concise.
JimDabell · · focus · HN ↗
I’ve noticed Astra doing this a lot as well.
runeks · · focus · HN ↗
GPT does this all the time, too (both Sol and Astra). I constantly have to tell it to not use terms that were not part of the initial prompt.
glub · · focus · HN ↗
ls612 · · focus · HN ↗
imron · · focus · HN ↗
Hah, yeah even when you put it in AGENTS.md or a skill.. constantly having to remind it.. "what does AGENTS.md" say about doing that?".. Thinking.. Thinking.. "Oh, it says I should never do that, I'll remember that next time.."
Next session - same thing.
suburban_strike · · focus · HN ↗
joegibbs · · focus · HN ↗
Rover222 · · focus · HN ↗
laurels-marts · · focus · HN ↗
So yea, I find fable 5.1 writing to be excellent everywhere. I still use Sol daily though, but for things like config, quick research, fixes, code review etc. Feature work and writing is for fable 5.1.
mrtesthah · · focus · HN ↗
I would pay money not to use Grok.
DoesntMatter22 · · focus · HN ↗
throw10920 · · focus · HN ↗
mrtesthah · · focus · HN ↗
aditya-ramabadr · · focus · HN ↗
Imustaskforhelp · · focus · HN ↗
That’s a factor of half the parameters. I would be curious to see more on the focus of smaller parameters model and pushing its frontiers
[deleted] · · focus · HN ↗
[deleted]
bushbaba · · focus · HN ↗
ndesaulniers · · focus · HN ↗
Lol, probably because Tesla's software stack is buildroot based. I'll bet that was in the training data.
octoberfranklin · · focus · HN ↗
And the fact that Grok is the ultimate grandmaster of parallel tool-calls, routinely kicking off four or five at once. Overlapping the latencies makes a huge difference in responsiveness.
I also like how Grok is trained to print a short one-sentence descriptions of what it's doing before each step. Like an airline pilot calling out observations for the black-box recorder to hear.
MaxQuimby · · focus · HN ↗
[dead]
kristofferR · · focus · HN ↗
Iolaum · · focus · HN ↗
babelfish · · focus · HN ↗
Jcampuzano2 · · focus · HN ↗
Its because of this. You can't use Astra in Cursor, and cursorbench uses cursor as the harness. They can't actually benchmark it using their harness hence why its not included.
babelfish · · focus · HN ↗
Jcampuzano2 · · focus · HN ↗
ryeguy · · focus · HN ↗
bluecalm · · focus · HN ↗
<a href="https://x.com/elonmusk/status/2102082011233931762?s=20" rel="nofollow">https://x.com/elonmusk/status/2102082011233931762?s=20
so it's likely about usage in Cursor specifically.
DavCreator · · focus · HN ↗
Jcampuzano2 · · focus · HN ↗
This explains why. Mentioned in another comment, but cursorbench explicitly tests with Cursor as the harness, and OpenAI doesn't allow them to use Astra in Cursor.
kristofferR · · focus · HN ↗
andsoitis · · focus · HN ↗
kristofferR · · focus · HN ↗
Jcampuzano2 · · focus · HN ↗
mh- · · focus · HN ↗
user43928 · · focus · HN ↗
That said, I don't expect them to benchmark Astra in their Cursor harness given the situation.
Jcampuzano2 · · focus · HN ↗
oh_no · · focus · HN ↗
scottyah · · focus · HN ↗
kristofferR · · focus · HN ↗
scottyah · · focus · HN ↗
toader · · focus · HN ↗
[dead]
AtlanticThird · · focus · HN ↗
thoman23 · · focus · HN ↗
ctrlkctrls · · focus · HN ↗
thereitgoes456 · · focus · HN ↗
Even with his successes (Tesla, SpaceX) he has built them up in large part by bending levers of government to his advantage.
voidfunc · · focus · HN ↗
So what? Thats called being a maverick. He is very very good at executing on making money which is the point of business.
andsoitis · · focus · HN ↗
Also pushing technology forward.
redox99 · · focus · HN ↗
blisterpeanuts · · focus · HN ↗
redox99 · · focus · HN ↗
Anthropic: 1.25B/month
Google: 0.92B/month
Unnamed customer starting in december: 1.1B/month
Starlink monthly revenue is ~1.5B/month
chris_money202 · · focus · HN ↗
redox99 · · focus · HN ↗
chris_money202 · · focus · HN ↗
sssilver · · focus · HN ↗
Can you provide specific examples of where Elon has bent the levers of government?
brandonagr2 · · focus · HN ↗
nozzlegear · · focus · HN ↗
keeda · · focus · HN ↗
chris_money202 · · focus · HN ↗
zamalek · · focus · HN ↗
toader · · focus · HN ↗
oulipo · · focus · HN ↗
jackfischer · · focus · HN ↗
estearum · · focus · HN ↗
If anything, they voted for reduced debt burden and they got the opposite. DOGE failed at pretty much every single one of the goals that the public arguably gave it a mandate for.
serbuvlad · · focus · HN ↗
Ah, yes, democracy!, except for when the public is wrong.
Who decides when the public is wrong? We do! Who decides "what the public voted for"? We do! So we are the rulers? No, of course, not, this is democracy.
You want to become the decider of when the public is wrong and of what the public voted for? TYRANT! TYRANT!
verdverm · · focus · HN ↗
half of voters don't pay any attention to politics until the week or two before voting
estearum · · focus · HN ↗
This is simply epistemologically incorrect. It's obviously incorrect in this case because voters writ large do not have any idea how the government is administered and how to improve it, so even if they claimed to be voting for that, it would not necessarily be an endorsement of any particular approach.
More specifically we know it's not true in this case because there are polls. Voters didn't even claim to care about this! "How the government is administered" was not a high salience issue to voters. Simple as that.
Nonetheless, I didn't suggest anything about overriding their votes. It sounds like you have some sensitive spots to work through (someone obliquely criticizing your idol for sucking at his job?)
maelito · · focus · HN ↗
fourseventy · · focus · HN ↗
[dead]
KyleTheDev · · focus · HN ↗
Sheep often like to think themselves the wolf or coyote, it would seem.
jml78 · · focus · HN ↗
Fuck, it is like the denial around Jan 6th. Those idiots we’re live streaming that shit. I watched it go down live. Now they say they weren’t violent.
We can’t have discourse when we have legit video evidence and people refuse to open their eyes and choose to deny reality
sejje · · focus · HN ↗
Which Nazi ideologies do you think he embraces? How do you reconcile all the Nazi ideologies he rejects?
oulipo · · focus · HN ↗
nancyminusone · · focus · HN ↗
butlike · · focus · HN ↗
"You're on, bro"
nibbleyou · · focus · HN ↗
ls612 · · focus · HN ↗
Romanulus · · focus · HN ↗
[dead]
jmward01 · · focus · HN ↗
andsoitis · · focus · HN ↗
jmward01 · · focus · HN ↗
solid_fuel · · focus · HN ↗
sejje · · focus · HN ↗
sidgtm · · focus · HN ↗
guywithahat · · focus · HN ↗
I'm excited for 4.7 although I share skepticism with other users whether 4.7 will be significantly better, since they didn't raise the price.
aschobel · · focus · HN ↗
vessenes · · focus · HN ↗
mrtesthah · · focus · HN ↗
[dead]
howunfortunate · · focus · HN ↗
I'm curious if you feel the same about re-migration of Belgians from the Congo?
Personally I think it's fine for any country to vote to control immigration as they see fit. I think Japan is a good example of a relatively xenophobic culture that deals with this fairly and thoughtfully.
thinkcontext · · focus · HN ↗
Can't say I've ever heard anyone implying that colonialists leaving Belgium was unjust. Colonialists is actually not the right word, more like extended occupation, only slightly better than the enslavement of the Leopold II era. The Belgian's were less than 1% of the population and all but an ancillary amount worked in exploiting the native population.
CuriousRose · · focus · HN ↗
bigyabai · · focus · HN ↗
throw10920 · · focus · HN ↗
Irrelevant. HN is not the place to randomly inject flamewars about politics. It's explicitly against both the purpose and guidelines of HN.
Seems like you need to review the guidelines again, because they're pretty clear:
> Eschew flamebait. Avoid generic tangents. Omit internet tropes.
<a href="https://news.ycombinator.com/newsguidelines.html">https://news.ycombinator.com/newsguidelines.html
bigyabai · · focus · HN ↗
The fact that Elon Musk's companies take contracts from the CIA and NRO is not flamebait. It's the lifeline of his company, strictly speaking.
UltraSane · · focus · HN ↗
Basic morality is not "flamewars" or "politics".
throw10920 · · focus · HN ↗
is not "basic morality" - it's entirely false and/or an unprovable claim (and therefore extremely obviously can't be about morality), but was posted because the author was trying to incite political flamewars on HN, which again, is against the guidelines.
You should review them, because you clearly don't know what's in them.
<a href="https://news.ycombinator.com/newsguidelines.html">https://news.ycombinator.com/newsguidelines.html
UltraSane · · focus · HN ↗
spiderice · · focus · HN ↗
We aren't responsible to fix your willful ignorance
UltraSane · · focus · HN ↗
thinkcontext · · focus · HN ↗
throw10920 · · focus · HN ↗
hardbass · · focus · HN ↗
mrtesthah · · focus · HN ↗
runsWphotons · · focus · HN ↗
mrtesthah · · focus · HN ↗
bigyabai · · focus · HN ↗
Tsarp · · focus · HN ↗
rvz · · focus · HN ↗
[0] <a href="https://quiver.ai/" rel="nofollow">https://quiver.ai/
jcims · · focus · HN ↗
kridsdale3 · · focus · HN ↗
user43928 · · focus · HN ↗
TylerE · · focus · HN ↗
lumirth · · focus · HN ↗
user43928 · · focus · HN ↗
That it isn't the most efficient way to achieve the same end result is irrelevant.
forgot-my-pw · · focus · HN ↗
Saline9515 · · focus · HN ↗
It will loop in thinking mode ("Let me implement those fixes: Fix 1, Fix 2, Fix 3 .... Fix 80, Fix 81"), ignore the AGENTS.md instructions, corrupt plan files, etc etc... I have 5.6 Sol as advisor/watchdog, and it blocks every turn, I never saw this. Quite a shame, 4.6 wasn't so bad.
xmorse · · focus · HN ↗
Saline9515 · · focus · HN ↗
samtheprogram · · focus · HN ↗
marwatk · · focus · HN ↗
I find it's very hard to get information on harnesses people are using. I have to stay model agnostic so I avoid claude, codex, cursor, etc. I've used and tried opencode, which worked well, but obviously lacks the above features.
Does anyone have a resource for following what people are actually being productive with? With so much vibe going on it's hard to separate the wheat from the chaff.
codeIMperfect · · focus · HN ↗
xmorse · · focus · HN ↗
<a href="https://x.com/greg_horvay/status/2100764473392820433?s=20" rel="nofollow">https://x.com/greg_horvay/status/2100764473392820433?s=20
raincole · · focus · HN ↗
Saline9515 · · focus · HN ↗
unrvl22 · · focus · HN ↗
polytely · · focus · HN ↗
raincole · · focus · HN ↗
andsoitis · · focus · HN ↗
totallymike · · focus · HN ↗
AM1010101 · · focus · HN ↗
ssutch3 · · focus · HN ↗
forgot-my-pw · · focus · HN ↗
ssutch3 · · focus · HN ↗
yearolinuxdsktp · · focus · HN ↗
I agree with GP, it's strange the comparison is not against Grok 4.6 xhigh.
everfrustrated · · focus · HN ↗
maz1b · · focus · HN ↗
avazhi · · focus · HN ↗
There for awhile it seemed like we’d have 3 big competitors but then Grok 4.2 or 4.4 was just diabolical while OAI and Claude continued their significant improvements. Grok was/is so bad that I was convinced musk was gonna shut it down and just fund Anthropic compute once they reached their compute agreement.
6thbit · · focus · HN ↗
asdfsa32 · · focus · HN ↗
But honestly, it is because numbers are like people; torture them enough and they'll tell you anything.
meerita · · focus · HN ↗
parineum · · focus · HN ↗
meerita · · focus · HN ↗
includenotfound · · focus · HN ↗
meerita · · focus · HN ↗
Well, at least I spent lots of dollars, and I had to use those models the same way I am using local and cheap models, with the same results.
includenotfound · · focus · HN ↗
On the other hand, Grok and GPT finish these tasks in <5min with no issues, and significantly better output.
meerita · · focus · HN ↗
includenotfound · · focus · HN ↗
testfrequency · · focus · HN ↗
user43928 · · focus · HN ↗
sparkling · · focus · HN ↗
RussianCow · · focus · HN ↗
But I can't argue with the lower off-peak pricing when using DeepSeek directly. The downside is they train their models on your input, which might be a deal-breaker for many users (as it is for me).
brcmthrowaway · · focus · HN ↗
microsoftedging · · focus · HN ↗
"Privacy# All these models are hosted in the US. Providers follow a zero-retention policy and do not use your data for model training, with the following exceptions:
Big Pickle: During its free period, collected data may be used to improve the model.
DeepSeek V4 Flash Free: During its free period, collected data may be used to improve the model.
MiMo-V2.5 Free: During its free period, collected data may be used to improve the model.
Laguna S 2.1 Free: During its free period, collected data may be used to improve the model.
Ling-3.0-tiny Free: During its free period, collected data may be used to improve the model.
LongCat-2.0 Free: During its free period, collected data may be used to improve the model.
North Mini Code Free: During its free period, collected data may be retained and used to improve the model. Do not submit personal or confidential data. See the provider’s Terms of Use and Privacy Policy.
Nemotron 3 Ultra Free (NVIDIA free endpoints): Trial use only — do not submit personal or confidential data. Your use is logged for security purposes and to improve NVIDIA products and services. The logged session data for improvement purposes is not linked to your identity or any persistent identifier. For more information about data processing practices, see the Privacy Policy. By interacting with this endpoint, you consent to the collection, recording, and use of such information and the NVIDIA API Trial Terms of Service."
BeetleB · · focus · HN ↗
thehamkercat · · focus · HN ↗
<a href="https://openrouter.ai/deepseek/deepseek-v4.1-flash?endpoint=57e1cb3e-9762-4ee1-a67c-259bd3af24b7#providers" rel="nofollow">https://openrouter.ai/deepseek/deepseek-v4.1-flash?endpoint=...
nicce · · focus · HN ↗
thehamkercat · · focus · HN ↗
riquito · · focus · HN ↗
thehamkercat · · focus · HN ↗
simlevesque · · focus · HN ↗
drewnick · · focus · HN ↗
_s_a_m_ · · focus · HN ↗
yipinwong · · focus · HN ↗
GLM or Kimi are better for my own personal projects. DS? uhm. it just keeps doing dumb crap
vorticalbox · · focus · HN ↗
In cursor I have switch over to grok for planning a composer for coding.
brianwawok · · focus · HN ↗
vorticalbox · · focus · HN ↗
thefourthchime · · focus · HN ↗
zug_zug · · focus · HN ↗
I can't think of a single dimension grok is winning on (capability, cost, voice), but want to stay open-minded -- anybody want to vouch for its capabilities in any domain?
grim_io · · focus · HN ↗
puszczyk · · focus · HN ↗
The voice is the same AI slop as the others imho.
(This is about Grok 4.6, I didn't test 4.7 yet).
svachalek · · focus · HN ↗
puszczyk · · focus · HN ↗
[dead]
Shekelphile · · focus · HN ↗
Deepswe results show that grok 4.6 is more expensive per-task and consistently scores worse than: luna xhigh, glm 5.3, astra low, sol high/xhigh, opus 5 medium.
Grok also used almost 3x as many tokens/turns to complete tasks than all of those models (besides luna), so it takes way more time to complete a task.
There isn't much reason to use Grok at all, it's gotten better but it's still worse than every other player in the field, which shouldn't be a surprise considering until about a year ago they were just buying tokens from other providers and pretending it was their own model.
sejje · · focus · HN ↗
I think it's winning on UI for normies (grok bot) and they made some claims about being pareto SOTA (lowest cost per task completed) a while back with 4.6.
I find it to be a perfectly capable model for implementation (there are many in this class--deepseek flash, spark1.3, luna, etc). I find the usage to be very generous w/ supergrok. I find the model to be just fine for 90% of what I want to do, but I use a smarter model to plan complicated things.
zug_zug · · focus · HN ↗
I'm judging on benchmarks, and whether anybody or any company I know has ever suggested using it (not yet).
sejje · · focus · HN ↗
I don't personally make my judgements based on how many other people mention a thing, but if that gets your code written, by all means.
swalsh · · focus · HN ↗
bluepeter · · focus · HN ↗
[dead]
simonw · · focus · HN ↗
sejje · · focus · HN ↗
simonw · · focus · HN ↗
btian · · focus · HN ↗
MuffinFlavored · · focus · HN ↗
Is there a metric for like... time taken when comparing these two? I see score and cost.
If Fable5.1 can knock it out more quickly on low but Grok4.7 might take twice as long to stumble through a problem (and leave behind a bunch of yucky comments or un-needed extra unit tests), are they really comparable?
Or like... the "quality" of the solution? "It works" versus "it's unmaintainable/very messy/hacky".
WarmWash · · focus · HN ↗
[1]<a href="https://eebench.org/" rel="nofollow">https://eebench.org/
saejox · · focus · HN ↗
xAI missed its chance, Ball is on Anthropic's court.
enraged_camel · · focus · HN ↗
It's phenomenal at computer use and 3D stuff. I've been using it less and less for coding.
brink · · focus · HN ↗
haellsigh · · focus · HN ↗
jquery · · focus · HN ↗
adventured · · focus · HN ↗
sneezychl · · focus · HN ↗
Best to stick with a high end model + low effort, do a manual pass on high effort and fix the bugs you know are reachable.
mcintyre1994 · · focus · HN ↗
cowboylowrez · · focus · HN ↗
redox99 · · focus · HN ↗
jstummbillig · · focus · HN ↗
brianwawok · · focus · HN ↗
nl · · focus · HN ↗
Elon claimed Opus was 5T in April, and I think it's fairly likely this is accurate: <a href="https://x.com/elonmusk/status/2042123561666855235" rel="nofollow">https://x.com/elonmusk/status/2042123561666855235
01100011 · · focus · HN ↗
manmal · · focus · HN ↗
imposter · · focus · HN ↗
[dead]
Razengan · · focus · HN ↗
pac0 · · focus · HN ↗
manmal · · focus · HN ↗
Razengan · · focus · HN ↗
ModernMech · · focus · HN ↗
Something else is right.
epolanski · · focus · HN ↗
The two models are in completely different price tiers. Astra costs 5 times as much.
It seems like all you can judge about cars would be their maximum speed on an oval.
user43928 · · focus · HN ↗
Based on Artificial Analysis Cost per Task, Astra is about 2-3x cheaper than Fable 5.1 at Medium and Low.
Consequently Astra could be cheaper than Grok 4.7, depending on the task.
simonw · · focus · HN ↗
Here's reasoning level high: <a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F8ab126bda2b384264b3ad931e3ebb8b4" rel="nofollow">https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
For some reason reasoning effort low and medium used similar numbers of tokens, and xhigh used less than high. I think I need to try without OpenRouter in the middle.
UPDATE: I tried again with the xAI API directly: <a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F5a1819a2bd24bb642f38c4bc6733090f" rel="nofollow">https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - not a great deal of difference between reasoning levels, and this time xhigh and low used the same number of reasoning tokens for some reason.
For comparison here's a fresh run against Grok 4.6: <a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ffedc404b9aa8e6fca31d59c898dabba0" rel="nofollow">https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
MattDamonSpace · · focus · HN ↗
datsci_est_2015 · · focus · HN ↗
kiliancs · · focus · HN ↗
forgot-my-pw · · focus · HN ↗
TomGarden · · focus · HN ↗
Mashimo · · focus · HN ↗
athrowaway3z · · focus · HN ↗
It used to be a mess in various interesting ways. Now, almost every big release can draw something perfectly functional.
So the question - without a correct answer - given the prompt "Generate an SVG of a pelican riding a bicycle":
Does the user want the least lines of code to make it functional, or the best looking version?
wolttam · · focus · HN ↗
peder · · focus · HN ↗
paimapi · · focus · HN ↗
daveguy · · focus · HN ↗
simonw · · focus · HN ↗
nicolamanzini · · focus · HN ↗
Grok 4.7 generations: <a href="https://threejseval.com/models/grok-4.7-high" rel="nofollow">https://threejseval.com/models/grok-4.7-high
Also go vote on <a href="https://threejseval.com" rel="nofollow">https://threejseval.com so you can help evaluate how Grok and other model performs compared to each other!
danappelxx · · focus · HN ↗
iankp · · focus · HN ↗
mrtesthah · · focus · HN ↗
jfoster · · focus · HN ↗
Has a shadow
Better shaped beak
Leg position more realistic for bicycle riding
Better feathers
dom96 · · focus · HN ↗
1 - <a href="https://bench.killswitch-lang.org" rel="nofollow">https://bench.killswitch-lang.org
sejje · · focus · HN ↗
For now, I doubt anyone would notice your protest if you didn't announce it.
lirolero · · focus · HN ↗
[dead]
peder · · focus · HN ↗
mempko · · focus · HN ↗
ulfw · · focus · HN ↗
thefourthchime · · focus · HN ↗
OrangeMusic · · focus · HN ↗
hersko · · focus · HN ↗
Totally the same.
ulfw · · focus · HN ↗
But a lot of young people on the Zuckerberg networks are falling for the nazi propaganda nowadays too around the world
mempko · · focus · HN ↗
cbeach · · focus · HN ↗
If Elon hadn’t worked with Orange Man Bad, then the Left would still be in love with him for his massive former donations to the Democrat political machine, and his work against climate change.
The whole “he’s a nazi” accusation is banal, and people are seeing through it now. That’s why we’ve moved on.
purerandomness · · focus · HN ↗
WarmWash · · focus · HN ↗
Like the whole pizza parlor pedo basement thing, people will death grip stupid stuff because they are so desperate to manifest the worst possible image of those unaligned with them.
The problem is that it blows up in their face and just makes them look unreliable, dumb, and lost.
Musk has done so many objectively bad things that there is no need for people to dilute their reputation on fringe theories and interpretations. Pushing the nazi thing just gives Musk ammo that his detractors are so desperate that they need freeze frames and hidden context to make him look bad.
dom96 · · focus · HN ↗
If you want him to get the benefit of the doubt about his "hand gesture" then it would help if he wasn't promoting far-right parties all over Europe.
sejje · · focus · HN ↗
They're arresting people for *praying silently in their mind*!!
The far left is everything they accuse the far right of being.
mempko · · focus · HN ↗
cisc · · focus · HN ↗
Here is the actual video in which Musk performs two fascist salutes at a political rally:
<a href="https://www.youtube.com/watch?v=smQNNo2a9xc" rel="nofollow">https://www.youtube.com/watch?v=smQNNo2a9xc
It wasn't an accident. He did it deliberately. Twice. He enjoyed himself doing it.
There's no value in the mental gymnastics to invent excuses for his fascist salutes. The rational thing to do is to accept the objective reality of it.
mempko · · focus · HN ↗
What's silly is people like you defending the richest man in the world who cares nothing about you and your family. Why spend the effort?
niek_pas · · focus · HN ↗
xerlait · · focus · HN ↗
felixgallo · · focus · HN ↗
knicholes · · focus · HN ↗
inferniac · · focus · HN ↗
oulipo · · focus · HN ↗
Vaslo · · focus · HN ↗
eleventen · · focus · HN ↗
[dead]
drop_star · · focus · HN ↗
ForrestN · · focus · HN ↗
moolcool · · focus · HN ↗
mlindner · · focus · HN ↗
Also it's kinda hilarious how you think any money spent on Grok will go toward furthering climate change versus literally any other AI model that does the same thing. Grok at least seems to be more efficient than most models.
grokgrokgrok · · focus · HN ↗
TheOtherHobbes · · focus · HN ↗
Handing corporate code secrets to his AI model is... unusually trusting.
mlindner · · focus · HN ↗
And methane is a large percentage of all power production in the US. So again that also applies to all the other data centers. (And FWIW they've been winding down and shutting down the on site methane generators.)
And no corporate code was handed to AI models.
moomin · · focus · HN ↗
jesse_dot_id · · focus · HN ↗
vb-8448 · · focus · HN ↗
unsupp0rted · · focus · HN ↗
eleventen · · focus · HN ↗
Nobody both worked and spent their money to get Trump elected like Musk. 300 million to his 2024 campaign [1]. DOGE. On-stage endorsements. Nobody even came close.
No, other big labs are not "innocent little virgins", but they're not even in the same solar system of harm as Musk. To hand-wave at the differences is to permit them.
[1] <a href="https://www.opensecrets.org/2024-presidential-race/donald-trump/contributors?id=N00023864" rel="nofollow">https://www.opensecrets.org/2024-presidential-race/donald-tr...
big_toast · · focus · HN ↗
These comments don't stay up much anymore and I can't tell if it's structural to the forum (flag weight + statistical mechanics of votes + guidelines) or if it's the userbase sentiment.
oulipo · · focus · HN ↗
eleventen · · focus · HN ↗
But I think it represents real malaise in the community. It's not a moderator plot, people here really just don't care and might even support this.
We really are in the minority of opinion for giving a damn about liberal democracy.
big_toast · · focus · HN ↗
Between the guidelines + user thoughts (e.g. repetition, low novelty/new info), there's other reasons these types of replies might end up dead.
I am worried that it leads to people self selecting to other forums biasing the remaining userbase vote distributions. In an exit vs voice situation, the voice kinda dies out.
hdhdjdif · · focus · HN ↗
DavideNL · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
TylerJaacks · · focus · HN ↗
[dead]
786562354238 · · focus · HN ↗
TuxSH · · focus · HN ↗
mrtesthah · · focus · HN ↗
oulipo · · focus · HN ↗
mavamaarten · · focus · HN ↗
swalsh · · focus · HN ↗
sparkling · · focus · HN ↗
Astra for deep dive investigations, Sol 5.6 at mid-level for day to day tasks, Grok 4.6 via Cursor for routine and low complexity tasks.
becquerel · · focus · HN ↗
johnfahey · · focus · HN ↗
stiltzkin · · focus · HN ↗
[dead]
shdtabasum · · focus · HN ↗
xquce · · focus · HN ↗
MoreThanMe · · focus · HN ↗
[dead]
thih9 · · focus · HN ↗
But also Xai doesn’t seem to care about user experience and long term support.
swozey · · focus · HN ↗
As a technical point of reference to compare against other llm stuff, sure, I'll glance at a report or benchmark but I really couldn't care less about anything to do with the project and it could blow other options away and I wouldn't touch it.
ElectronCharge · · focus · HN ↗
You probably shouldn't cut off your nose to spite your face.
mempko · · focus · HN ↗
What's superficial about refusing to use a product from someone like that? Or are you one of those 'technology isn't about politics' people? That's a superficial take if you ask me.
All technology is political, and understanding that is a deep, not superficial take. It requires systems thinking which unfortunately many people building technology seem to lack, despite software being a sophisticated complex system.
totallymike · · focus · HN ↗
brandonagr2 · · focus · HN ↗
venzaspa · · focus · HN ↗
eknkc · · focus · HN ↗
For daily one off questions I prefer it because it is fast enough and I like the way it responds. I also use it for basic research like “find me a battery drill for this and that”.
Kimi and GLM feel extremely coding oriented. I hate the way Anthropic models talk. GPT takes too much time and effort for that kind of stuff for some reason.
Grok happened to be a nice middle ground.
jpadkins · · focus · HN ↗
nimchimpsky · · focus · HN ↗
[dead]
thih9 · · focus · HN ↗
> somehow this high profile AI model seems more disgusting than others and it is in a way impressive.
thefourthchime · · focus · HN ↗
gslepak · · focus · HN ↗
andreyvit · · focus · HN ↗
nwienert · · focus · HN ↗
everfrustrated · · focus · HN ↗
For me and what I’m doing that’s insanely good value.
I find grok build chews through my SuperGrok sub very quick - but I think that is due to it having the 500k context window which uses more credits. Cursor limits it to 256K (tho I see in today’s update for Grok 4.7 there’s now a toggle for context size).
thefourthchime · · focus · HN ↗
daquisu · · focus · HN ↗
It is the same multiplier for Sol with subscription. For Astra though the multiplier is 20x, so half of Sol usage.
For Claude it seems to be around 40x too for Opus, but less for Fable (similar to Astra in GPT).
Some sources:
1. <a href="https://x.com/kunchenguid/status/2098256018836963382" rel="nofollow">https://x.com/kunchenguid/status/2098256018836963382
2. <a href="https://x.com/stevenzhang/status/2092110386569089311" rel="nofollow">https://x.com/stevenzhang/status/2092110386569089311
3. <a href="https://github.com/openai/codex/issues/43731" rel="nofollow">https://github.com/openai/codex/issues/43731
4. <a href="https://redd.it/1wciwc1" rel="nofollow">https://redd.it/1wciwc1
5. <a href="https://x.com/SemiAnalysis_/status/2064815044085318040" rel="nofollow">https://x.com/SemiAnalysis_/status/2064815044085318040
6. <a href="https://redd.it/1vx0k69" rel="nofollow">https://redd.it/1vx0k69
nwienert · · focus · HN ↗
daquisu · · focus · HN ↗
I can't help much more than that, I did that research in the last few days, but I never used Grok myself.
I pay for GPT, Claude and Gemini. Last week I consumed all my quota on two of them, so I wondered which next subscription I would pay for if needed.
nwienert · · focus · HN ↗
revenue_sub_doe · · focus · HN ↗
revenue_sub_doe · · focus · HN ↗
artdigital · · focus · HN ↗
Normal SuperGrok barely lasts me through the week with very mild usage and no coding. The sentiment around SuperGrok Plus is also not great and I haven’t seen someone saying they’re happy with it yet.
SuperGrok Heavy is $300/mo, so you could get a full ChatGPT Pro and Claude Max 5x for that price. That’s so far out of my budget for a single provider I haven’t bothered trying it.
I still have SuperGrok through X Premium+ but will downgrade that next billing cycle
notduckrabbit · · focus · HN ↗
everfrustrated · · focus · HN ↗
notduckrabbit · · focus · HN ↗
sourcecodeplz · · focus · HN ↗
- grok 4.6 (xhigh): 97M (for 44 score)
- grok 4.7 (xhigh): 240M (for 46 score)
c0rruptbytes · · focus · HN ↗
mh- · · focus · HN ↗
GodelNumbering · · focus · HN ↗
From their headline comparison:
Long running agentic workflows are dominated by cache reads.Just makes Grok sound deceptive, and more importantly, reliant on user's lack of understanding of costs aka predatory (which in turn is more infuriating)
sourcecodeplz · · focus · HN ↗
bastawhiz · · focus · HN ↗
qwerpy · · focus · HN ↗
Excited to try 4.7. I hope they fixed the "it's not X, it's Y" that showed up in 4.6.
Theodores · · focus · HN ↗
By now AI should know of the DRY concept. But no. Hence the keys have a rounded rectangle for the key shape and another rounded rectangle for a clip path, to prevent text overflow. There are 72 * 2 = 144 identical rectangles, when just one would suffice (in the defs), with this being cloned once for the clip path, and 72 times for the keys.
I would not expect SVGO levels of optimisation (rounding numbers, that sort of thing), however, the human, if writing out the same thing for the 72nd time, might think 'is there a better way', to get the manual out. A graphics program such as Illustrator would not do that, but AI 'should' because AI.
The above is not criticism of your work, just an observation regarding AI SVG capabilities.
qwerpy · · focus · HN ↗
Theodores · · focus · HN ↗
There are also interesting inheritance rules with SVG, so you could define the basic shape of a key, well, several shapes, just as rects in the defs, with no stroke or fill specified.
Then, at the group level, you can then specify stroke and fill, so there could be a group of normal keys, another group for modifiers, function keys and so on.
Then there are the keys themselves, how do you clone a shape and put different text inside each clone? There are many ways to do this but I think you are on the right track using the clip path approach, albeit using the rects in the defs.
What is interesting about SVG is that artists don't care for the file format, they just see text as shapes on a page. Then programmers don't care for SVG as that is a graphic designer/artworker thing. So SVG sits in this witch-space, with only a few brave enough to wade in and do cool stuff.
Given your application, and given the fun that could be had with SMIL/JS, you could make your SVG files interactive, so you press a key and a popover tells you more about what that key does. You can even get audio working in SVG, as well as HTML popovers (in foreignobjects, as buttons, but working, nonetheless).
'Views' is another interesting SVG feature. I have a sprite sheet that uses a lot of views, where you are projecting your SVG into some type of virtual canvas, taking a 'picture' of it, and then incorporating that in something else, maybe a CSS variable.
One 'deadly addiction' is animation. Filters are another 'deadly addiction'. Why have a static and actually useful diagram, when you can animate it, move the 'camera' and add the equivalent of 27 Photoshop layers as filters to everything?
For example, supposing you wanted to show what keys to press, with there being modifiers and a sequence, e.g. 'Hello World!'. The animation for one letter could be what triggers the animation for the next 'key press' and so it goes.
My top tip of all: reposition the origin (0,0) to where it makes sense. Many objects have symmetry, so you can define one side, clone it, scale it (-1,1) and do it all around 0,0 to then translate the results to somewhere sensible.
I have found the JetBrains IDEs to be extremely useful for SVG, the preview feature is very helpful, as are the code hints.
AI sort of knows SVG, so I have had some suggestions from Google on how to build filters. These never work, but they do get you thinking. Say you wanted to use filters to add specular highlights and animated shadows to the keys, that would be fair game for AI hints on how to do it.
mrtesthah · · focus · HN ↗
qwerpy · · focus · HN ↗
enraged_camel · · focus · HN ↗
trentor · · focus · HN ↗
brrrrrm · · focus · HN ↗
BoumTAC · · focus · HN ↗
<a href="https://x.com/ValsAI/status/2102086608476590432" rel="nofollow">https://x.com/ValsAI/status/2102086608476590432
nostrebored · · focus · HN ↗
BoumTAC · · focus · HN ↗
I like to follow them and look for benchmark for each LLM release.
nostrebored · · focus · HN ↗
brcmthrowaway · · focus · HN ↗
hdhdjdif · · focus · HN ↗
musk can fund the space stuff with this
dgellow · · focus · HN ↗
sourcecodeplz · · focus · HN ↗
Output tokens from Intelligence Index:
- grok 4.6 (xhigh): 97M (for 44 score)
- grok 4.7 (xhigh): 240M (for 46 score)
oh_no · · focus · HN ↗
gaigalas · · focus · HN ↗
usumgallu · · focus · HN ↗
[dead]
oh_no · · focus · HN ↗