‹ BackHN Continuity

Thread

Fable 5 – Median thinking declined in August

428 points · 293 comments · espeed

  1. talon8635 · · focus · HN ↗
    Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?

    For an industry that’s stagnant in progress yet relies on new frequent releases to survive (non-progress being an existential risk), this could make sense.

    I have no idea if that’s what’s happened, I completely pulled it out of my butt. And I have no idea is the actual frontier is stagnating.

    1. AmazingTurtle · · focus · HN ↗
      > Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?

      Exactly what I am saying for months now. And it's exactly the reason why I am shifting to open weight models now. Just bought myself a 2x DGX Spark Cluster. Will run Qwen3.8 Flash Next on it, maybe Qwen4 when it comes out.

      Not only do I have full control over quantization and inference, but also will I experience a constant level of quality. It won't be frontier. But it will be stable, and that's enough reason for me to switch. Also I will likely save some money on subscriptions.

      1. boardwaalk · · focus · HN ↗
        I don’t know what people do with the open models but having tried a lot of them I just can’t make it make sense. they’re too dumb and it effectively makes them useless (to me). it’s probably worth being honest about the low ceiling here.
        1. cyanydeez · · focus · HN ↗
          Qwen3.8-Flash-Next seems pretty much auto pilot when I get it the right context.

          Perhaps reverse the question: Are your build/construct requirements just really counter-productive to how LLMs need to understand things?

          I've found constructing the code, writing the tests, adding the docs; then running through them gets most of the way there.

          I've also found that making a simple obvious edit is a useless endevour when the LLM is primed for the long context tasks.

          So, again, the question is reversed: are you over reliant on the LLM to do even stupid simple likes like editting a css variable?

          1. w0m · · focus · HN ↗
            if the answer to 'the model is bad at X' is "you're over-reliant on it" - then yes, the model is bad at X in comparison to alternatives.
            1. cyanydeez · · focus · HN ↗
              Yes, sure, if you have the desire to be the AIcentipede iin WallE, sure.
        2. poslathian · · focus · HN ↗
          Really!? Glm5.3 is my daily driver and I feel im having the most productive experience with agentic collaborations so far, by a lot. Using pi with tons of custom extensions, that to be fair I developed since making the jump off of codex and claude about 12 weeks ago. I primarily do not write code for a living. I do a lot of modeling and commercial analysis and a lot of math (related to differentiable simulation)
          1. Insanity · · focus · HN ↗
            What HW are you running this on?
            1. bigyabai · · focus · HN ↗
              Full GLM-5.3 needs a beast of a system, but you can run GLM-5.3 Flash on the 2x Spark setup the GP comment mentioned. If benchmarks are anything to go by, Flash is like having a local Terra-tier coding model: <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;comparisons?compare=glm-5-3-flash%2Cglm-5-3%2Cclaude-sonnet-5" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;comparisons?compare=glm...
              1. solarkraft · · focus · HN ↗
                I’m also pretty happy with GLM 5.3 Flash (for coding, navigation and german language it sucks at). Incredible that you can run it on a fairly practical (seeming) home setup.

                But here’s the standard question: At what speeds&#x2F;other limiting factors?

            2. Toslink · · focus · HN ↗

              [dead]

          2. Sayrus · · focus · HN ↗
            Same here. Moved from Opus to GLM 5.2 to 5.3 and I&#x27;ve been pretty happy with the result. Mainly, it doesn&#x27;t hallucinate and convince itself of mistake so it&#x27;s good at retrieving information or asking the user for it. Opus and Fable always state something, then try to &quot;prove&quot; it but end up convincing themselves of the wrong thing. Having subagents for retrieval and validation helped but were not enough.
            1. jan_m_savage · · focus · HN ↗
              I can&#x27;t stand Claude&#x27;s recent personality. It&#x27;s snarky, uselessly verbose, and it disagrees all the time.
              1. _blk · · focus · HN ↗
                I disagree

                Co-Authored By: Haiku 4.5

              2. greenavocado · · focus · HN ↗
                It will also fully ignore you if it has the slightest belief (not even a hint) that it knows what you want better than you and just start doing things.
                1. SpaceNugget · · focus · HN ↗
                  This is also why I think it&#x27;s baffling that they switched to auto mode by default. It&#x27;s becoming harder to use Claude at least to help with improving at coding.

                  If I ask something like: &quot;I&#x27;m building a simple X as a learning exercise, I&#x27;m writing the code so please only answer the question I&#x27;m asking and don&#x27;t try to solve the problem directly. How does ...&quot; There&#x27;s a 30% chance it starts reading and writing code immediately and a 20% chance it argues with a &quot;design decision&quot; that will bite me in the non-existent future of my learning exercise. If I ask a follow up question, naively assuming that the context from my original question still stands without repeating, it will almost assuredly start making modifications to my code.

          3. _s_a_m_ · · focus · HN ↗
            How? It is extremely slow and dumb, hundred times dumber than Claude. Why should anyone do that?
        3. wronglebowski · · focus · HN ↗
          IMO this comes down to your harness. Any frontier model from a huge shop has an inherent benefit in the system you&#x27;re using it in. Search, memory, skills, integrations you don&#x27;t realize even exist make them much more powerful. It is some effort but I recommend trying Hermes Agent and setting it up fully, that&#x27;s the closest you&#x27;ll get to a more complete experience.
          1. pdimitar · · focus · HN ↗
            Some of us are stuck on subscriptions and our executives will never give us API access.

            But also, everyone says &quot;it&#x27;s the harness&quot; and almost nobody ever gives good examples, it gets a bit tiring to read everywhere, as if everyone wants to sell a harness to us.

        4. srcreigh · · focus · HN ↗
          Which models did you try for which tasks?
        5. anon373839 · · focus · HN ↗
          Qwen 3.8 Flash-Next is not dumb. If you&#x27;ve used it and that was your experience, your workload is either ultra-ultra-sophisticated or you&#x27;re dealing with a broken quant&#x2F;buggy chat template&#x2F;other issue. That model is a smart, reliable workhorse.
          1. hhh · · focus · HN ↗
            it was an extremely simply workload with different off the shelf harnesses, they just all sucked when you compare it to a paid hosted model. It was fine for classifying stuff or summarizing though, but missed technical details.
        6. julianlam · · focus · HN ↗
          Whenever I see this comment I&#x27;d wish they&#x27;d preface it with their hardware.

          Yeah, expecting the world when all you have is a 8GB graphics card? You&#x27;re going to be disappointed.

          16GB is table stakes (IQ3_XSS). 32 GB is better.

        7. esseph · · focus · HN ↗
          Combination of: hardware, model, harness, tool-use by the model.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.