‹ BackHN Continuity

Thread

Prompting Claude Opus 5.5

207 points · 227 comments · Michelangelo11

  1. Aissen · · focus · HN ↗
    > test several levels against your own evals

    Of course, and this is the basics anyone should do when working with LLMs & agents; but with their high-variance, doing statistically significant benchmarking is very costly. Which is why the debates here on HN often talk about the "feelings" of degradation (or improvement!), but often without proofs. I'm not sure how to solve ạt; maybe inference providers should provide free benchmarking to anyone publishing results, along with the guarantee to never train on those sessions.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.