‹ BackHN Continuity

Thread

One month coding with GLM 5.3 Flash

228 points · 180 comments · ThibWeb

  1. fallinditch · · focus · HN ↗
    > For day-to-day developer work, it’s totally viable to focus on one or two flash-tier cheap models

    GLM 5.3 flash has been good as my default profile Hermes bot, after I readjusted its memory to point to a couple of key skills.

    I got some great coding results with GLM 5.2, and 5.3 Flash is supposedly almost as good, so I will be trying it out soon for day to day tasks as the post advises.

    1. UncleOxidant · · focus · HN ↗
      I've been using GLM 5.3 Flash quite a lot for coding this month since they've had their promotion running. It's been tackling some difficult stuff - coding up a Julia version of the luminal GPU kernel optimization package, FPGA work with Verilog (including getting a BitNet 2B model running on FPGA - that one's been split between Claude and GLM), coding up a Julia version of DiffLUT. It's been handling these tasks pretty well. I do move between Claude Sonnet 5.5 & GLM 5.3-flash on the BitNet one based on what's available.
      1. buildbot · · focus · HN ↗
        > BitNet 2B model running on FPGA Very cool! What kind of FPGA & What kind of TPS are you hitting?
        1. UncleOxidant · · focus · HN ↗
          It's still in process so we're not getting tokens out of it yet. Targetting a Gowin 138K (the Tang Retro Console 138K board). This is moving from a Xilinx part that's about 2x that size so we have to trade some time for space. Estimated to be in the 7tok/sec when it's all said and done. Not super fast, but this isn't a super fast (clockrate) FPGA. And at 138K LUTs it's pretty small for this kind of task (but it is the largest FPGA that I've got, and the boards are quite affordable at $129 which includes 8GB of DDR3 - which is another reason it's kind'a of slow, DDR3 is kind'a slow)
      2. [deleted] · · focus · HN ↗

        [deleted]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.