‹ BackHN Continuity

Thread

Beating GPT-5.6 Sol on retrieval with 100x cheaper open models

415 points · 114 comments · moonikakiss

  1. oedemis · · focus · HN ↗
    whats about self-consistency like in grpo with majority voiting?
    1. kumama · · focus · HN ↗
      castform founder here. i'm personally a little against techniques like self-consistency/majority voting during rl training because they tend to result in the model's output distribution "sharpening" a lot. this means the model will lose it's exploration ability and probably won't be able to explore/discover new solution strategies, which can be harmful for both rl training + generalization to unseen cases
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.