‹ BackHN Continuity

Thread

Fable 5 – Median thinking declined in August

428 points · 293 comments · espeed

  1. jotato · · focus · HN ↗
    Just yesterday I was thinking about gpt-5.6-luna. I made it my default model in Hermes during its fist week of launch. It was just as good as 5.5 which was my previous default. But over the last 2 or 3 weeks I've seen how dumb it is now. I have to be very explicit with it.

    For example, I used to be able to prompt "Check the system logs on <server> for...." and it would just figure it out. Yesterday I asked "Did <service> on <server> complete the overnight job" and all it said was "that service is not installed on my host"

    I had to tell it to ssh into the server and run journlctl to check it

    Anecdotal, I know, but they all seem to be less capable with time.

    _edit_ I use the same reasoning level of `medium`

    1. ajspig1 · · focus · HN ↗
      & the nice thing about Hermes (since its open source) is you can be reasonably sure that behavior change is coming from the model and not the harness. (probably)
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.