How do you think other benchmarks work? Every single one is going to have a budget and time limit. There has to be a cap on those, you can't just let it spin forever and use unlimited funds.
I would argue that this benchmark is uniquely suited to how most people use LLMs because it actually tests common harnesses and it is a true coding benchmark for a language that is unseen, thus testing the LLMs actual ability to understand nuance and learn.
Internet access is restricted to ensure that over time models cannot just look up the source code of KillSwitch, which would allow them to cheat.
I didn't expect there to be a time limit and I do know benchmarks that don't have a time limit. They just penalize slower models which I think is fair.
In my experience models just don't take forever to mark tasks as done.
For my coding usage I don't set time limit so benchmarks that do so are providing me less interesting use cases.
With that said, I can understand having budget limits for expensive LLMs for those of us that don't have infinite VC money.
As for no internet access, the issue is that for most problems we do want Llms to be able to search docs, API specs, GitHub issues, etc. So not allowing that usually just favours larger LLMs that were able to memorize more data, not necessarily smarter ones when both have internet access.
dom96 · · focus · HN ↗
KillSwitch-Bench 1.0
1 - <a href="https://bench.killswitch-lang.org/" rel="nofollow">https://bench.killswitch-lang.org/bel8 · · focus · HN ↗
- capped per-task budget and time limit
- No internet access
- different harnesses mixed
dom96 · · focus · HN ↗
I would argue that this benchmark is uniquely suited to how most people use LLMs because it actually tests common harnesses and it is a true coding benchmark for a language that is unseen, thus testing the LLMs actual ability to understand nuance and learn.
Internet access is restricted to ensure that over time models cannot just look up the source code of KillSwitch, which would allow them to cheat.
bel8 · · focus · HN ↗
In my experience models just don't take forever to mark tasks as done.
For my coding usage I don't set time limit so benchmarks that do so are providing me less interesting use cases.
With that said, I can understand having budget limits for expensive LLMs for those of us that don't have infinite VC money.
As for no internet access, the issue is that for most problems we do want Llms to be able to search docs, API specs, GitHub issues, etc. So not allowing that usually just favours larger LLMs that were able to memorize more data, not necessarily smarter ones when both have internet access.