It's also the only model that generates accurate translation and localization. No other frontier model comes close. Although Gemini's coding capabilities are subpar, its natural language processing is top-tier.
I’m curious how you guys keep track of each model’s coding capabilities. The landscape keeps changing. I don’t suppose you benchmark all frontier models every other month, right?
Do you dangerously allow permissions? I absolutely cannot use it until they ship an auto approver. As it is now I have it write one bash/python script to do everything it wants to, then I review that. Otherwise it is COMPLETELY unusable and it shocks me when I hear people are using it.
Zsfe510asG · · focus · HN ↗
phenomen · · focus · HN ↗
thisgoodlife · · focus · HN ↗
riddlemethat · · focus · HN ↗
safog · · focus · HN ↗
Most of the time I don't need what the bench tests and I'm not really giving them completely ambiguous tasks without any refinement.
I only find marginal differences between models at this point and it almost feels like personality quirks in each model than anything.
BenzeneDream · · focus · HN ↗
vrosas · · focus · HN ↗
taylorfinley · · focus · HN ↗
cute_boi · · focus · HN ↗