The actual tokens might be non-deterministic, but you could look for proxy measures that are supposed to be invariant. Eg. correctness/performance on benchmarks, "thinking level" on complex problems, etc
This is a good overview of how this is done: <a href="https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents" rel="nofollow">https://www.anthropic.com/engineering/demystifying-evals-for...
cloudking · · focus · HN ↗
ssivark · · focus · HN ↗
CharlesW · · focus · HN ↗
6gvONxR4sf7o · · focus · HN ↗