On the smaller end, Quen 3.8, while being extraordinarily capable for a small local model, also suffers from extreme thinking. I wonder if the techniques described here generalize to other models too.
<a href="https://www.reddit.com/r/LocalLLaMA/comments/1wj3s31/thank_you_swift_qwen_38_27b_now_has_100k/" rel="nofollow">https://www.reddit.com/r/LocalLLaMA/comments/1wj3s31/thank_y...
ls612 · · focus · HN ↗
KaoruAoiShiho · · focus · HN ↗