‹ BackHN Continuity

Thread

DeepSeek v4.1 Flash Is Now Our Best Hacking Model

177 points · 67 comments · talhof8

  1. TuxSH · · focus · HN ↗
    I find this - or perhaps the title - a bit surprising.

    I've benchmarked GLM 5.3 and DSv4.1-F on my fully-annotated decomp of the Nintendo 3DS's kernel, which I have a good mental understanding of, tasking them to find vulns and other bugs (in Max mode w/ subagents). GLM 5.3 founds almost all the vulns in 30min for $22, while DS only found one vuln for $2 in 40min.

    Perhaps DS works better where targets have low-hanging fruits than can be found fast?

    1. mariopt · · focus · HN ↗
      Makes sense, GLM 5.1/5.2/5.3 is a lot better than DS v4.1, I found the same results in other domains.

      Given how fast and cheap DS is, it's just an ideal model with enough "IQ" to let it loose. Another thing they left out of the article, DS becomes really good if you provide custom tools for the task, on it's own it's mediocre.

      1. seemaze · · focus · HN ↗
        How does GLM 5.3 Flash rank against it's big brother and the latest Deepseek?
        1. networked · · focus · HN ↗
          Worth saying that GLM-5.3 isn&#x27;t GLM-5.3-Flash&#x27;s &quot;big brother&quot; the way one might think. GLM-5.3-Flash is not GLM-5.3 scaled down. While GLM-5.3 is based on GLM-5.2, and &quot;every gain comes from post-training&quot; (<a href="https:&#x2F;&#x2F;z.ai&#x2F;blog&#x2F;glm-5.3" rel="nofollow">https:&#x2F;&#x2F;z.ai&#x2F;blog&#x2F;glm-5.3), GLM-5.3-Flash uses a newly trained multimodal base model (<a href="https:&#x2F;&#x2F;z.ai&#x2F;blog&#x2F;glm-5.3-flash" rel="nofollow">https:&#x2F;&#x2F;z.ai&#x2F;blog&#x2F;glm-5.3-flash).
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.