Clef: Open-weight decision models, and new RL fine-tuning platform
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Clef: Open-weight decision models, and new RL fine-tuning platform
Unofficial Hacker News client; not affiliated with Y Combinator.
manlymuppet · · focus · HN ↗
And it's only been a few weeks.
segmondy · · focus · HN ↗
SebastianSosa · · focus · HN ↗
Foobar8568 · · focus · HN ↗
verdverm · · focus · HN ↗
- I'd like a potion, are you sure, no, repeat
- in and out of doors on loop
- sisyphean effort in the cave
- jev-ish level grinding
It was impressive, beat pokemon for less than $2, but not all that interesting. People asking how different Math.Random plays pokemon would be, and at the other end, regular llms playing games.
indoor47 · · focus · HN ↗
"Clef builds upon this concept, but uses a different base model as the backbone. We currently use Qwen as the base model and post-trained it to suit decision model use cases. "