GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Unofficial Hacker News client; not affiliated with Y Combinator.
Aboutplants · · focus · HN ↗
At the same time, OpenAI is also making its existing $200 Pro plan less appealing. In Codex and Work, $200 Pro subscribers will see their included usage decrease from 20x of what the company offers to Plus users, down to 10x of that same allowance. In ChatGPT, meanwhile, GPT-6 Pro message caps will decrease from 200 to 100 per week.”
<a href="https://www.engadget.com/2272106/openai-adds-dollar500-pro-subscription-nerfs-its-existing-dollar200-tier/" rel="nofollow">https://www.engadget.com/2272106/openai-adds-dollar500-pro-s...
Yikes
latentsea · · focus · HN ↗
tripleee · · focus · HN ↗
latentsea · · focus · HN ↗
That they are expensive and climbing doesn't negate my point if the cost of the subscription over how long you plan to keep it is equally or more expensive than the GPUs. You can put together dual 5060 Ti or 5070 Ti systems to run local LLMs too. You don't need to splurge on a 5090. That's a bad option at this point.
tripleee · · focus · HN ↗
I've messed around with Qwen3.6-27B but I'm not sure if it could yet even replace Luna for me.
latentsea · · focus · HN ↗
Qwen3.8-Flash-Next is better still if you can run fast enough. If you have a dual R9700 setup you certainly can. That model is even better.
Qwen4-27B has been announced but not released yet. I'm super pumped for it because I already use 3.8 as my daily driver at home for all my personal stuff, so I'm definitely happy to take an increase in capability.
There is clearly still room for improvement in local models on consumer hardware. With the Qwen 27B models, If you have at least a 5070 Ti I think you can get away with running a small Q4 quant if you use KV cache streaming. The 24GB cards can run Q4 comfortably. If you have a 32B card you can run Q6 comfortably. If you have 48GB ~ 64GB of VRAM you can Q8 comfortably. Using llama.cpp Vulkan let's you pool VRAM across cards (even AMD and NVIDIA etc), so my machine has a 5060 Ti and an R9700.
A dual R9700 rig is really the sweet spot right now with the vLLM-radiance fork. If you can swing a 5070 Ti in there as well to retain some CUDA access, then all the better. That's basically the equivalent to spending 2 years on a subscription, but gets you a system that can run Qwen-3.8-Flash-Next and of course the even more capable Qwen4-Flash when it releases. At the end of the two years it'll run even better models I'm sure.
I'm all in on local now.
directdev · · focus · HN ↗
Would you be up for a 20-minute chat about your experience, or a few lines by email?