> For day-to-day developer work, it’s totally viable to focus on one or two flash-tier cheap models
GLM 5.3 flash has been good as my default profile Hermes bot, after I readjusted its memory to point to a couple of key skills.
I got some great coding results with GLM 5.2, and 5.3 Flash is supposedly almost as good, so I will be trying it out soon for day to day tasks as the post advises.
I've been using GLM 5.3 Flash quite a lot for coding this month since they've had their promotion running. It's been tackling some difficult stuff - coding up a Julia version of the luminal GPU kernel optimization package, FPGA work with Verilog (including getting a BitNet 2B model running on FPGA - that one's been split between Claude and GLM), coding up a Julia version of DiffLUT. It's been handling these tasks pretty well. I do move between Claude Sonnet 5.5 & GLM 5.3-flash on the BitNet one based on what's available.
It's still in process so we're not getting tokens out of it yet. Targetting a Gowin 138K (the Tang Retro Console 138K board). This is moving from a Xilinx part that's about 2x that size so we have to trade some time for space. Estimated to be in the 7tok/sec when it's all said and done. Not super fast, but this isn't a super fast (clockrate) FPGA. And at 138K LUTs it's pretty small for this kind of task (but it is the largest FPGA that I've got, and the boards are quite affordable at $129 which includes 8GB of DDR3 - which is another reason it's kind'a of slow, DDR3 is kind'a slow)
fallinditch · · focus · HN ↗
GLM 5.3 flash has been good as my default profile Hermes bot, after I readjusted its memory to point to a couple of key skills.
I got some great coding results with GLM 5.2, and 5.3 Flash is supposedly almost as good, so I will be trying it out soon for day to day tasks as the post advises.
UncleOxidant · · focus · HN ↗
buildbot · · focus · HN ↗
UncleOxidant · · focus · HN ↗