Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)
Unofficial Hacker News client; not affiliated with Y Combinator.
breckenedge · · focus · HN ↗
tedsanders · · focus · HN ↗
GPT-5.6 Sol's performance in the API should not change over time. If it has, that's a severe bug and we'll look into it.
We do sometimes tweak ChatGPT settings (e.g., tools, system prompts, efforts) over time, but we never play games to juice evals at launch times. You should always get what's advertised.
(I work at OpenAI.)
sunaurus · · focus · HN ↗
billypilgrim · · focus · HN ↗
tedsanders · · focus · HN ↗
Serial testing over time is much less reliable than side by side testing, and even when I do side by side testing, I try to look at multiple attempts per prompt. Seeing multiple per prompt helps me realize how much intrinsic variation there is. My brain always wants to see patterns even when there isn’t enough data to prove them.