‹ BackHN Continuity

Thread

Several vulnerabilities have been discovered in the Linux kernel

576 points · 408 comments · luispa

  1. boutell · · focus · HN ↗
    I head a much, much smaller open source project. Since the November Singularity we've been seeing at least six responsibly reported security advisories a month. However, this last month we had 22 unique security advisories. Our project has been built with adherence to the OWASP Top Ten Guidelines and other best practices from the beginning. But software is hard and AI is thorough.

    Each month, we fix them all in our monthly maintenance release and disclose at that time. We fight AI fire with fire, and hand-review, of course.

    So far, we can keep up. One hopes this is possible at the scale of the Linux project, which assuredly has more humans and more AI to throw at the problem. But team size does not scale linearly with interested audience, and potential bugs do scale with codebase size (and other extremely important factors, like code quality, at which the Linux team is assuredly much better than we are).

    ("November Singularity" is a cheeky reference to the arrival of Opus 4.5 and "good enough" coding models and harnesses generally.)

    1. Roark66 · · focus · HN ↗
      I wonder if August of this year will be remembered like that too. The very first "opus like" local AI model came out this August (Qwen 3.8 Flash Next). I've been running it locally since for real programming and I consider it pretty much the same as opus 4.6 in coding ability (it lacks a bit in the factual knowledge area). It even exceeds opus on some tasks.

      I've tried every hyped "open source" model before and all including latest models bigger than 1T parameters are pretty much toys.

      This is the first one that isn't. It can't be overstated how huge of a deal that is. No more reliance on Anthropic.

      Running this model to do real work is still not cheap. I run it on a pc with 6 rtx3090s and 192GB of RAM (and I use 90gb of that ram for KV cache). The model is entirely on gpus. It runs at 55tok/s 1500 prefill for one user at a time, and about 35tok/s 650 prefill per user for 5 simultaneous users. It doesn't seem like much until one realises you manage your own kv cache. You ca leave your sessions in cache for as long as you want. You can save them and restore 200k sessions a week later in a dozen seconds.

      What many people don't realise is that usage of those models skews extremely heavily towards input processing. My claude code usage is about 1.3B tokens input per week and only about 8M output. On claude code I get 80% cache. At home it's more like 95%.

      If this progress keeps up, and we get a fable quality model in a year in under 200B to run at home... Those "frontier labs" will be renting all of their gpus per hour not to go bankrupt.

      1. QwenGlazer9000 · · focus · HN ↗
        What Quants do you run?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.