‹ BackHN Continuity

Thread

DeepSeek v4.1 Flash Is Now Our Best Hacking Model

177 points · 67 comments · talhof8

  1. TuxSH · · focus · HN ↗
    I find this - or perhaps the title - a bit surprising.

    I've benchmarked GLM 5.3 and DSv4.1-F on my fully-annotated decomp of the Nintendo 3DS's kernel, which I have a good mental understanding of, tasking them to find vulns and other bugs (in Max mode w/ subagents). GLM 5.3 founds almost all the vulns in 30min for $22, while DS only found one vuln for $2 in 40min.

    Perhaps DS works better where targets have low-hanging fruits than can be found fast?

    1. severino · · focus · HN ↗
      A little off-topic: where does one use those models such as GLM or DS for this kind of reverse engineering tasks? I think I read many of them refuse to help with tasks like those on their official platforms.
      1. TuxSH · · focus · HN ↗
        GLM 5.3 doesn't seem to refuse vuln research (which it classifies as "audit") and is good at it.

        Therefore you use for offensive cybersecurity tasks because Daybreak Red/Mythos is pure unobtainium for us mere plebians.

        DB Blue thankfully exists, but I suspect you risk a ban if you use it with codebases you neither own nor use

        Tl;dr because it's the only model at the level of 5.4~5.6 that doesn't refuse tasks nor risk your oai account getting banned

        For plain RE tasks Sol or Astra should work just fine (I think)

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.