‹ BackHN Continuity

Thread

Ember-1

589 points · 249 comments · gmays

  1. soerxpso · · focus · HN ↗
    I don't see why I would be interested in this model, considering the price difference. They advertise that it's the same as Kimi K3 in half the tokens. But the pricing is double the pricing of K3. So why do I care if it uses fewer tokens, if I'm paying double per token?
    1. bigmadshoe · · focus · HN ↗
      Because you care about how many tokens are used per task. What you said is like only caring about the price of gas and not gas mileage of your car.
      1. verdverm · · focus · HN ↗
        it's not exactly the same, the model stills "weighs" the same

        here, it does less work, it's more like driving half as far but still paying the same total cost

        this being said, K3 and E1 models are priced the same at $3.00 / $0.30 / $15.00

        <a href="https:&#x2F;&#x2F;fireworks.ai&#x2F;models&#x2F;fireworks&#x2F;kimi-k3" rel="nofollow">https:&#x2F;&#x2F;fireworks.ai&#x2F;models&#x2F;fireworks&#x2F;kimi-k3

        <a href="https:&#x2F;&#x2F;fireworks.ai&#x2F;models&#x2F;fireworks&#x2F;ember-1" rel="nofollow">https:&#x2F;&#x2F;fireworks.ai&#x2F;models&#x2F;fireworks&#x2F;ember-1

        1. bigmadshoe · · focus · HN ↗
          From reading the blog post, it is essentially exactly the same as the car example. It delivers the same performance on tasks, but using 40% fewer tokens. This is the same as a car getting you to the same destination but wasting less energy on excess heat, wind resistance, or whatever else affects fuel economy (I am not an expert, obviously). I am paying for an LLM to complete tasks for me, not for the intermediate tokens.
          1. verdverm · · focus · HN ↗
            I&#x27;m with you, I wrote prior comment under the assumption that they were priced differently (from GP claim as such), but they are priced the same (on Fireworks)

            Will be taking Ember-1 for a spin on Monday and hopefully enjoy those better MPGs

        2. ranguna · · focus · HN ↗
          Why are you saying it drives half as far when both models complete the same task but one uses half the tokens?

          Wouldn&#x27;t the op be more correct with their gas&#x2F;distance comparison?

          Because both cars get to the same end destination (complete the same task).

          Unless your end goal is to see the token numbers go up, but I&#x27;m not sure why that would be of interest.

    2. ersiees · · focus · HN ↗
      It’s same price for lower latency.
    3. verdverm · · focus · HN ↗
      the pricing is the same, where are you seeing double?
      1. tyingq · · focus · HN ↗
        K3 is cheaper from other providers. $1&#x2F;$9, though I can&#x27;t speak to how good those providers are.

        <a href="https:&#x2F;&#x2F;openrouter.ai&#x2F;moonshotai&#x2F;kimi-k3" rel="nofollow">https:&#x2F;&#x2F;openrouter.ai&#x2F;moonshotai&#x2F;kimi-k3

        1. verdverm · · focus · HN ↗
          as the saying goes, you get what you pay for

          we require ZDR and Fireworks provides that on contract, so for us they are the same price

          1. indigodaddy · · focus · HN ↗
            Check out Neuralwatt. They are ZDR and great energy based K3 pricing. They also have a K3-fast which is basically no reasoning (in addition to regular K3).
            1. verdverm · · focus · HN ↗
              that is nothing like how we consume Ai

              industry standard is price-per-1M tokens, don&#x27;t do something different, even Google caved and moved from their char based pricing to tokens (the fundamental unit of computation in ai)

              GPUs are rented in $&#x2F;h, like every other piece of hardware in cloud

              1. Bolwin · · focus · HN ↗
                Why are you assuming the gpus are rented?

                Anyway, they still have token based pricing if you prefer. The energy pricing is often cheaper though

                1. verdverm · · focus · HN ↗
                  Not interested, it looks amateur

                  A trust page is required like <a href="https:&#x2F;&#x2F;trust.fireworks.ai" rel="nofollow">https:&#x2F;&#x2F;trust.fireworks.ai

                  Very happy with my $10&#x2F;month OpenCode Go sub for personal use

                  <a href="https:&#x2F;&#x2F;trust.opencode.ai&#x2F;">https:&#x2F;&#x2F;trust.opencode.ai&#x2F;

              2. indigodaddy · · focus · HN ↗
                Energy based is way more fundamental than token based. To each their own.
        2. polski-g · · focus · HN ↗
          K3 TOS says they must sell no lower than what Moonshot charges. If you have a provider selling for less than $15&#x2F;mtok, they are violating TOS from Moonshot.
          1. verdverm · · focus · HN ↗
            looks like the underlying vendor is charging correct prices <a href="https:&#x2F;&#x2F;inference.net&#x2F;models&#x2F;kimi-k3&#x2F;" rel="nofollow">https:&#x2F;&#x2F;inference.net&#x2F;models&#x2F;kimi-k3&#x2F;

            however the listing on open router has a `&#x2F;fp4` suffix, so perhaps this is an unlisted, quanted model for a lower price?

    4. seizethecheese · · focus · HN ↗
      Half the tokens presumably means tasks get done twice as fast.
    5. zupa-hu · · focus · HN ↗
      Agentic work costs ~input^2. For a single message, they cost the same. For long conversations, you end up paying much less.
    6. XCSme · · focus · HN ↗
      I mean, Gpt-6 Luna seems a lot better in every aspect:

      <a href="https:&#x2F;&#x2F;aibenchy.com&#x2F;compare&#x2F;fireworks-ember-1-high&#x2F;openai-gpt-6-luna-high&#x2F;" rel="nofollow">https:&#x2F;&#x2F;aibenchy.com&#x2F;compare&#x2F;fireworks-ember-1-high&#x2F;openai-g...

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.