Very capable model. I just ran it on my own LLM benchmark suite[1] and it matches Muse Spark 1.3 in pass rate but is significantly cheaper.KillSwitch-Bench 1.0 Claude Opus 5 66.9 GPT-6 Astra 57.9 Claude Fable 5.1 46.7 MiMo-V2.6-Pro 38.8 Muse Spark 1.3 36.5 1 - <a href="https://bench.killswitch-lang.org/" rel="nofollow">https://bench.killswitch-lang.org/
I appreciate that this benchmark is different but it is in no way how most people use LLMs or promote equal grounds when benchmarking:- capped per-task budget and time limit- No internet access- different harnesses mixed
dom96 · · focus · HN ↗
KillSwitch-Bench 1.0
1 - <a href="https://bench.killswitch-lang.org/" rel="nofollow">https://bench.killswitch-lang.org/bel8 · · focus · HN ↗
- capped per-task budget and time limit
- No internet access
- different harnesses mixed
yt1998 · · focus · HN ↗
[dead]