Several vulnerabilities have been discovered in the Linux kernel
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Several vulnerabilities have been discovered in the Linux kernel
Unofficial Hacker News client; not affiliated with Y Combinator.
boutell · · focus · HN ↗
Each month, we fix them all in our monthly maintenance release and disclose at that time. We fight AI fire with fire, and hand-review, of course.
So far, we can keep up. One hopes this is possible at the scale of the Linux project, which assuredly has more humans and more AI to throw at the problem. But team size does not scale linearly with interested audience, and potential bugs do scale with codebase size (and other extremely important factors, like code quality, at which the Linux team is assuredly much better than we are).
("November Singularity" is a cheeky reference to the arrival of Opus 4.5 and "good enough" coding models and harnesses generally.)
sparklingmango · · focus · HN ↗
jitl · · focus · HN ↗
kridsdale1 · · focus · HN ↗
soulofmischief · · focus · HN ↗
gamerdonkey · · focus · HN ↗
Why isn't even easier for an AI noob to jump in at the next step with less friction than the current one?
slopinthebag · · focus · HN ↗
soulofmischief · · focus · HN ↗
The person you're replying to had the decency to ask me for clarification; for you, I'd suggest brushing up on your reading comprehension and learning to respect the rules of this site, which include interpreting comments in their best light and focusing on positive, substantial contributions instead of negative, inflammatory posts.
slopinthebag · · focus · HN ↗
soulofmischief · · focus · HN ↗
tyg13 · · focus · HN ↗
Part of this is certainly that my hand-authoring code skills have atrophied, sure, but my workflow has also radically changed. Previously, I would spent a lot of time and focus on a single work item, and only context switch to other tasks whenever I would wait on CI or a long build. It meant that I spent a lot of time understanding one thing at a time, and interruptions (forced context switches) incurred a massive switching cost.
Now, having moved to largely AI-authored code, I find myself necessarily working on multiple threads at the same time. This means I can meaningfully progress each of those threads in parallel, with a much-reduced overhead on context switching, since I don't have my head down focusing on all the details of the work. And if the task really demands it, I can still stop and focus on one thread to sketch out the code manually, think about the concepts more deeply, etc.
It's a very different workflow, and there are certainly downsides, but the upside is that the rate of which I've been able to put up good-quality PRs has measurably increased. It's not quite 2x, and it's certainly not 10x, but it's definitely noticeable. I do understand a bit less, but there was always more work than time to understand things in full detail. I guess only time will tell if that missing understanding was actually vital to the long-term success of my work.
slopinthebag · · focus · HN ↗
becquerel · · focus · HN ↗
sysguest · · focus · HN ↗
hmm doesn't mean noob will become better? if what I'm "learning" about AI is going to get deprecated every week... then it doesn't make sense to "keep up"
andybak · · focus · HN ↗
1. Both are multipliers on skills you already have. A "noob" will be at a disadvantage.
2. The second one is a complex blend of technical skill, domain knowledge and human factors (knowing your user-base, knowing your product/project, understanding UX/DX/whatever) that you probably will always have an edge on.
Roark66 · · focus · HN ↗
Also knowing what models are good for what. There was time half a year ago when Google genuinely had a better model than everyone else. I used it for everything and 3 weeks later they lobotomise it (sorry "optimised") and I went back to opus...
I think Anthropic is like a drug dealer, giving us the sweet sweet drug for free (I don't think our $200 a month subscriptions even cover the electricity for our use) and the time to pay will come very soon...
I expect this subscription will cost $2k a month. Will it be a normal increase? Or will the "enshittify" existing models to the point you'll pay $2k to get fable 6 to do what opus 4.8 did fine in July 2026?
soulofmischief · · focus · HN ↗
What will that differential be? A large component will be social reinforcement: generational wealth, connectedness, and such. Things like merit might take more of a backseat. So people who do not have the necessary social capital may be fighting each other for very limited amounts of positions.
That is the steepening of the curve for people like me: I was homeless at 16 and finished high school on my own, was given a full ride to LSU, grants, room and board and a job in the Comp Sci department, but lost all of it after an immature and vindictive high school teacher illegally modified my grade in a core class in order to fuck me over. I didn't have parents to back me up at the school board and make things right.
Instead I suffered through years of homelessness and had to find my own path into the industry by starting companies with friends and doing all of the engineering. Since then I've led multiple teams, made some connections, shipped a lot of cool stuff and bring to the table a wealth of experience and a generalist skillset that is both wide and deep. Yet, I too wonder where my place in this changing industry will be once things settle a bit. Probably less engineering and more focus on business development.
The flipside though is as you've said: The fruits of engineering are more accessible than ever to the layman, and individuals can currently possess an unprecedented amount of agency and leverage. I think that is amazing and am fully behind it. I do know that it means the process of renormalization is going to be very rough, given the similarly unprecedented rate of industry change these technologies are bringing.
sandworm101 · · focus · HN ↗
It is like people fishing. Some people go out and catch fish with the tools they know will work. For other people, every day at the lake requires a new boat/rod/lure. They spend more time figuring out how to use their new toy than they do catching fish.
eli_gottlieb · · focus · HN ↗
ptidhomme · · focus · HN ↗
Secret knowledge/data will be tomorrow's gold.
gewetensleegte · · focus · HN ↗
that sounds ominous. what do you mean?
Germanioum · · focus · HN ↗
jrflo · · focus · HN ↗
ACS_Solver · · focus · HN ↗
November 2025, with Opus 4.5, was the first time I was impressed by an LLM doing something non-trivial with a reasonably good level of quality.
boutell · · focus · HN ↗
KiwiJohnno · · focus · HN ↗
boutell · · focus · HN ↗
phkahler · · focus · HN ↗
Wouldn't it be nice if AI vulnerability reports came with AI pull requests to fix them? The thinking context that found it should be readily able to propose a fix. It would still need review but even when AI PRs aren't right they often point in the right direction.
trklausss · · focus · HN ↗
This just goes to show that yes, if security researchers were to do that, it would great, but they are not the only actors here...
boutell · · focus · HN ↗
And in most cases they do propose solutions, although we generally resolve the issues on our own.
There's a small percentage where we make the case that the ticket is not a real vulnerability, and then we have to grit our teeth through repeated reports of the same "vulnerability." But it's a small percentage so far.
We do typically have to reconsider the severity. The researchers understandably want to see everything as a nine...
fsmv · · focus · HN ↗
0xdeadbeefbabe · · focus · HN ↗
eli · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
iririririr · · focus · HN ↗
boutell · · focus · HN ↗
tosapple · · focus · HN ↗
[dead]
Roark66 · · focus · HN ↗
I've tried every hyped "open source" model before and all including latest models bigger than 1T parameters are pretty much toys.
This is the first one that isn't. It can't be overstated how huge of a deal that is. No more reliance on Anthropic.
Running this model to do real work is still not cheap. I run it on a pc with 6 rtx3090s and 192GB of RAM (and I use 90gb of that ram for KV cache). The model is entirely on gpus. It runs at 55tok/s 1500 prefill for one user at a time, and about 35tok/s 650 prefill per user for 5 simultaneous users. It doesn't seem like much until one realises you manage your own kv cache. You ca leave your sessions in cache for as long as you want. You can save them and restore 200k sessions a week later in a dozen seconds.
What many people don't realise is that usage of those models skews extremely heavily towards input processing. My claude code usage is about 1.3B tokens input per week and only about 8M output. On claude code I get 80% cache. At home it's more like 95%.
If this progress keeps up, and we get a fable quality model in a year in under 200B to run at home... Those "frontier labs" will be renting all of their gpus per hour not to go bankrupt.
QwenGlazer9000 · · focus · HN ↗
roosterIllusi0n · · focus · HN ↗
qwen3.8 is more than capable for software dev. The gemma models were terrible and fell apart during compaction. I think as long as the llm can test the results, you don't need anything being offered by a "frontier" model. The cloud AI is going to be used by people unable to run their own and they will eventually get squeezed on price.
The one tip I could give is compact before starting new steps or any action in the plan that is different than what was previously worked on. You want to manage what is in the context and don't want unnecessary work history details filling it up. You can always ask the llm to list the current plan, then compact after and do it more than once until you get the compaction <30%. This will be fixable by the harness that can choose better times to compact.
You don't want to start a new phase and have it compact a few minute after starting. This happening over and over again seems to potentially cause issues for long running sessions. Compaction slop that screws up what is in the context.
I can compare what I use at home vs paid models like astra at work and the difference is mostly meaningless.
vlovich123 · · focus · HN ↗
lukan · · focus · HN ↗
And those who go bankrupt, can sell their GPU's cheap, so I can indeed finally have my own fable.
keeda · · focus · HN ↗
The thing was the wrangling was relatively straightforward, if cumbersome. Largely, it involved being very precise with the context and instructions it was given. I could imagine a lot of that getting automated (i.e. what we today call harnesses) or recursively addressed by creative meta-prompting. Supported by similarly conceptually simple advances like chain-of-thought reasoning I suspect that is the biggest thing that the models have figured out what to do today compared to then: manage themselves carefully.
Although I could not have predicted these exact outcomes, the implications for everything that is unfolding now were clear even then.
jasondigitized · · focus · HN ↗
qingcharles · · focus · HN ↗
Every single time I've done this it has found at least one serious bug or security hole.
Makes me wonder how many are left I've not found.
roosterIllusi0n · · focus · HN ↗
ozim · · focus · HN ↗
As much as cybersec forums were outraged that everyone will be hacked because of that — nothing happened.
aaroninsf · · focus · HN ↗
cookiengineer · · focus · HN ↗
Instead we should aim for quick updates, strong isolation and sandboxing.
Whatever that means for TDD and other methodologies that seemingly all have failed to encode guardrails in the development workflow.
fittingopposite · · focus · HN ↗
theteapot · · focus · HN ↗