‹ BackHN Continuity

Thread

Fable 5 – Median thinking declined in August

428 points · 293 comments · espeed

  1. jotato · · focus · HN ↗
    Just yesterday I was thinking about gpt-5.6-luna. I made it my default model in Hermes during its fist week of launch. It was just as good as 5.5 which was my previous default. But over the last 2 or 3 weeks I've seen how dumb it is now. I have to be very explicit with it.

    For example, I used to be able to prompt "Check the system logs on <server> for...." and it would just figure it out. Yesterday I asked "Did <service> on <server> complete the overnight job" and all it said was "that service is not installed on my host"

    I had to tell it to ssh into the server and run journlctl to check it

    Anecdotal, I know, but they all seem to be less capable with time.

    _edit_ I use the same reasoning level of `medium`

    1. Starlevel004 · · focus · HN ↗
      I'm fairly sure it's just luck of the draw if you get put onto a quant'd model or not. I've seen luna xhigh change intelligence fairly drastically on a day to day basis.
      1. pixl97 · · focus · HN ↗
        Really this is the base problem. You have zero idea where and how your prompt is being executed.

        If for example AWS sells you a 2xLarge server there may be some variability in performance but it's going to be averaged out very well.

        When it comes to AI services executing your model there is absolutely no information on what and with what settings your model is being executed. Hell, you have no idea if it even is the model you're paying for. Add that models are not deterministic so variability can be pretty large.

        This leads to a common set of dynamics that induce cheating behavior in humans. For example, is there a mix of different hardware. Does lessor hardware use different settings? How do you know xhigh is what your prompt ran under. Anthropic has a proven history of running your prompt silently under different models.

        This is a huge mess that needs and will be regulated or sued heavily over. Hell, with as many people out there that hate AI it might be easier than one thinks to have a state sue the providers on this and elicit a huge amount of discovery.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.