It'd be silly to buy the 18k model to run a tiny model like Qwen 27B. You use models like GLM Flash and Qwen Next which won't fit on a single 5090.
You might be. Running another agent doesn't load a set of new weights. It creates a new KV cache for the agent and adds the prompts to the queue. Its just another inference turn.
ApolloFortyNine · · focus · HN ↗
I didn't expect this to make the 5090 to look like a good deal.
nacs · · focus · HN ↗
It'd be silly to buy the 18k model to run a tiny model like Qwen 27B. You use models like GLM Flash and Qwen Next which won't fit on a single 5090.
asimovDev · · focus · HN ↗
Eisenstein · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]