‹ BackHN Continuity

Thread

DeepSeek v4.1 Flash Is Now Our Best Hacking Model

177 points · 67 comments · talhof8

  1. TuxSH · · focus · HN ↗
    I find this - or perhaps the title - a bit surprising.

    I've benchmarked GLM 5.3 and DSv4.1-F on my fully-annotated decomp of the Nintendo 3DS's kernel, which I have a good mental understanding of, tasking them to find vulns and other bugs (in Max mode w/ subagents). GLM 5.3 founds almost all the vulns in 30min for $22, while DS only found one vuln for $2 in 40min.

    Perhaps DS works better where targets have low-hanging fruits than can be found fast?

    1. simlevesque · · focus · HN ↗
      In your example, couldn't you parallelize DS's work more ? You could have 11 times as many agents for the same price.
      1. TuxSH · · focus · HN ↗
        Both DS and GLM had the same numbers of subagents, 5 or so.

        But, well, number of subagents doesn't make a difference if model is dumb (GPT 5.4 High, in May,in Chat mode outperformed what I see with DS4.1-F).

        That being said, pricing model makes a huge difference for "find at least one" tasks: with API/PAYG if you have a chance to save 90%, you go for it, whereas with subscriptions it is optimal to burn all your remaining allowance right before reset

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.