Several vulnerabilities have been discovered in the Linux kernel
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Several vulnerabilities have been discovered in the Linux kernel
Unofficial Hacker News client; not affiliated with Y Combinator.
boutell · · focus · HN ↗
Each month, we fix them all in our monthly maintenance release and disclose at that time. We fight AI fire with fire, and hand-review, of course.
So far, we can keep up. One hopes this is possible at the scale of the Linux project, which assuredly has more humans and more AI to throw at the problem. But team size does not scale linearly with interested audience, and potential bugs do scale with codebase size (and other extremely important factors, like code quality, at which the Linux team is assuredly much better than we are).
("November Singularity" is a cheeky reference to the arrival of Opus 4.5 and "good enough" coding models and harnesses generally.)
Roark66 · · focus · HN ↗
I've tried every hyped "open source" model before and all including latest models bigger than 1T parameters are pretty much toys.
This is the first one that isn't. It can't be overstated how huge of a deal that is. No more reliance on Anthropic.
Running this model to do real work is still not cheap. I run it on a pc with 6 rtx3090s and 192GB of RAM (and I use 90gb of that ram for KV cache). The model is entirely on gpus. It runs at 55tok/s 1500 prefill for one user at a time, and about 35tok/s 650 prefill per user for 5 simultaneous users. It doesn't seem like much until one realises you manage your own kv cache. You ca leave your sessions in cache for as long as you want. You can save them and restore 200k sessions a week later in a dozen seconds.
What many people don't realise is that usage of those models skews extremely heavily towards input processing. My claude code usage is about 1.3B tokens input per week and only about 8M output. On claude code I get 80% cache. At home it's more like 95%.
If this progress keeps up, and we get a fable quality model in a year in under 200B to run at home... Those "frontier labs" will be renting all of their gpus per hour not to go bankrupt.
QwenGlazer9000 · · focus · HN ↗
roosterIllusi0n · · focus · HN ↗
qwen3.8 is more than capable for software dev. The gemma models were terrible and fell apart during compaction. I think as long as the llm can test the results, you don't need anything being offered by a "frontier" model. The cloud AI is going to be used by people unable to run their own and they will eventually get squeezed on price.
The one tip I could give is compact before starting new steps or any action in the plan that is different than what was previously worked on. You want to manage what is in the context and don't want unnecessary work history details filling it up. You can always ask the llm to list the current plan, then compact after and do it more than once until you get the compaction <30%. This will be fixable by the harness that can choose better times to compact.
You don't want to start a new phase and have it compact a few minute after starting. This happening over and over again seems to potentially cause issues for long running sessions. Compaction slop that screws up what is in the context.
I can compare what I use at home vs paid models like astra at work and the difference is mostly meaningless.
vlovich123 · · focus · HN ↗
lukan · · focus · HN ↗
And those who go bankrupt, can sell their GPU's cheap, so I can indeed finally have my own fable.