I don't see why I would be interested in this model, considering the price difference. They advertise that it's the same as Kimi K3 in half the tokens. But the pricing is double the pricing of K3. So why do I care if it uses fewer tokens, if I'm paying double per token?
From reading the blog post, it is essentially exactly the same as the car example. It delivers the same performance on tasks, but using 40% fewer tokens. This is the same as a car getting you to the same destination but wasting less energy on excess heat, wind resistance, or whatever else affects fuel economy (I am not an expert, obviously). I am paying for an LLM to complete tasks for me, not for the intermediate tokens.
I'm with you, I wrote prior comment under the assumption that they were priced differently (from GP claim as such), but they are priced the same (on Fireworks)
Will be taking Ember-1 for a spin on Monday and hopefully enjoy those better MPGs
Check out Neuralwatt. They are ZDR and great energy based K3 pricing. They also have a K3-fast which is basically no reasoning (in addition to regular K3).
industry standard is price-per-1M tokens, don't do something different, even Google caved and moved from their char based pricing to tokens (the fundamental unit of computation in ai)
GPUs are rented in $/h, like every other piece of hardware in cloud
K3 TOS says they must sell no lower than what Moonshot charges. If you have a provider selling for less than $15/mtok, they are violating TOS from Moonshot.
looks like the underlying vendor is charging correct prices <a href="https://inference.net/models/kimi-k3/" rel="nofollow">https://inference.net/models/kimi-k3/
however the listing on open router has a `/fp4` suffix, so perhaps this is an unlisted, quanted model for a lower price?
soerxpso · · focus · HN ↗
bigmadshoe · · focus · HN ↗
verdverm · · focus · HN ↗
here, it does less work, it's more like driving half as far but still paying the same total cost
this being said, K3 and E1 models are priced the same at $3.00 / $0.30 / $15.00
<a href="https://fireworks.ai/models/fireworks/kimi-k3" rel="nofollow">https://fireworks.ai/models/fireworks/kimi-k3
<a href="https://fireworks.ai/models/fireworks/ember-1" rel="nofollow">https://fireworks.ai/models/fireworks/ember-1
bigmadshoe · · focus · HN ↗
verdverm · · focus · HN ↗
Will be taking Ember-1 for a spin on Monday and hopefully enjoy those better MPGs
ranguna · · focus · HN ↗
Wouldn't the op be more correct with their gas/distance comparison?
Because both cars get to the same end destination (complete the same task).
Unless your end goal is to see the token numbers go up, but I'm not sure why that would be of interest.
ersiees · · focus · HN ↗
verdverm · · focus · HN ↗
tyingq · · focus · HN ↗
<a href="https://openrouter.ai/moonshotai/kimi-k3" rel="nofollow">https://openrouter.ai/moonshotai/kimi-k3
verdverm · · focus · HN ↗
we require ZDR and Fireworks provides that on contract, so for us they are the same price
indigodaddy · · focus · HN ↗
verdverm · · focus · HN ↗
industry standard is price-per-1M tokens, don't do something different, even Google caved and moved from their char based pricing to tokens (the fundamental unit of computation in ai)
GPUs are rented in $/h, like every other piece of hardware in cloud
Bolwin · · focus · HN ↗
Anyway, they still have token based pricing if you prefer. The energy pricing is often cheaper though
verdverm · · focus · HN ↗
A trust page is required like <a href="https://trust.fireworks.ai" rel="nofollow">https://trust.fireworks.ai
Very happy with my $10/month OpenCode Go sub for personal use
<a href="https://trust.opencode.ai/">https://trust.opencode.ai/
indigodaddy · · focus · HN ↗
polski-g · · focus · HN ↗
verdverm · · focus · HN ↗
however the listing on open router has a `/fp4` suffix, so perhaps this is an unlisted, quanted model for a lower price?
seizethecheese · · focus · HN ↗
zupa-hu · · focus · HN ↗
XCSme · · focus · HN ↗
<a href="https://aibenchy.com/compare/fireworks-ember-1-high/openai-gpt-6-luna-high/" rel="nofollow">https://aibenchy.com/compare/fireworks-ember-1-high/openai-g...