It's also the only model that generates accurate translation and localization. No other frontier model comes close. Although Gemini's coding capabilities are subpar, its natural language processing is top-tier.
I’m curious how you guys keep track of each model’s coding capabilities. The landscape keeps changing. I don’t suppose you benchmark all frontier models every other month, right?
Do you dangerously allow permissions? I absolutely cannot use it until they ship an auto approver. As it is now I have it write one bash/python script to do everything it wants to, then I review that. Otherwise it is COMPLETELY unusable and it shocks me when I hear people are using it.
Sounds like they shipped some changes today that might reduce approvals: <a href="https://x.com/antigravity/status/2100001904969297980" rel="nofollow">https://x.com/antigravity/status/2100001904969297980
I've used Antigravity as my main coding agent on one of my biggest projects for about a year. It's been great for me. (and I use Claude, Codex, Grok and Muse for all the other projects)
I did a test involving implementing cobol control flow in Java for a source to source translation project. Gemini was the only model to get the edge cases. Cobol is very peculiar in this regard.
It's very good at Elixir in my experience too. And it just does what I ask and doesn't wind me up like Opus. I don't think I've had to insult it more than once per day.
I had a typical $20 Gemini plan that I just downgraded to their $5 plan (to keep access to some of the models). It had been so long since I let Gemini work on (or review) any code / design / html (anything) that I couldn't justify bothering to keep wasting money on it. It fell behind badly over the past year. Astra might as well be an alien super intelligence at code compared to Gemini. I enjoy talking to Gemini, it is very good at conversation, I get solid answers to everyday questions. I intend to keep the $5 plan indefinitely for basic use. I don't expect they'll ever resurface as a competitor in coding with Astra & Fable et al.
Zsfe510asG · · focus · HN ↗
phenomen · · focus · HN ↗
thisgoodlife · · focus · HN ↗
riddlemethat · · focus · HN ↗
baq · · focus · HN ↗
safog · · focus · HN ↗
Most of the time I don't need what the bench tests and I'm not really giving them completely ambiguous tasks without any refinement.
I only find marginal differences between models at this point and it almost feels like personality quirks in each model than anything.
BenzeneDream · · focus · HN ↗
vrosas · · focus · HN ↗
taylorfinley · · focus · HN ↗
xnx · · focus · HN ↗
andai · · focus · HN ↗
cute_boi · · focus · HN ↗
vrosas · · focus · HN ↗
akho · · focus · HN ↗
VectorLock · · focus · HN ↗
qingcharles · · focus · HN ↗
le-mark · · focus · HN ↗
robotmay · · focus · HN ↗
adventured · · focus · HN ↗