‹ BackHN Continuity

Thread

Fable 5 – Median thinking declined in August

428 points · 293 comments · espeed

  1. talon8635 · · focus · HN ↗
    Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?

    For an industry that’s stagnant in progress yet relies on new frequent releases to survive (non-progress being an existential risk), this could make sense.

    I have no idea if that’s what’s happened, I completely pulled it out of my butt. And I have no idea is the actual frontier is stagnating.

    1. AmazingTurtle · · focus · HN ↗
      > Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?

      Exactly what I am saying for months now. And it's exactly the reason why I am shifting to open weight models now. Just bought myself a 2x DGX Spark Cluster. Will run Qwen3.8 Flash Next on it, maybe Qwen4 when it comes out.

      Not only do I have full control over quantization and inference, but also will I experience a constant level of quality. It won't be frontier. But it will be stable, and that's enough reason for me to switch. Also I will likely save some money on subscriptions.

      1. boardwaalk · · focus · HN ↗
        I don’t know what people do with the open models but having tried a lot of them I just can’t make it make sense. they’re too dumb and it effectively makes them useless (to me). it’s probably worth being honest about the low ceiling here.
        1. cyanydeez · · focus · HN ↗
          Qwen3.8-Flash-Next seems pretty much auto pilot when I get it the right context.

          Perhaps reverse the question: Are your build/construct requirements just really counter-productive to how LLMs need to understand things?

          I've found constructing the code, writing the tests, adding the docs; then running through them gets most of the way there.

          I've also found that making a simple obvious edit is a useless endevour when the LLM is primed for the long context tasks.

          So, again, the question is reversed: are you over reliant on the LLM to do even stupid simple likes like editting a css variable?

          1. w0m · · focus · HN ↗
            if the answer to 'the model is bad at X' is "you're over-reliant on it" - then yes, the model is bad at X in comparison to alternatives.
            1. cyanydeez · · focus · HN ↗
              Yes, sure, if you have the desire to be the AIcentipede iin WallE, sure.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.