I think they even more so need deterministic feedback:
I tried an approach based on the following idea recently and it's amazing - Lint rules where the error messages contain an explanation on how to deal with the issue.
Deterministic feedback is precisely how frontier models are trained. It’s called RLVR. You let the agent run on a problem and then calculate a deterministic score of how well it did. Repeat 1000x times and you can “brute force” a good solution. (Which includes all thinking traces and you add it to your training data.) And then a Chinese model can copy your advance for 1000x less compute. Which is why US labs call this not learning, but a distillation “attack”. It’s an attack on the business model.
Labs typically pay $2k for each [prompt+scorer] docker image. So this is why frontier labs need so much cash and manpower and they see it as their moat.
Garlef · · focus · HN ↗
I tried an approach based on the following idea recently and it's amazing - Lint rules where the error messages contain an explanation on how to deal with the issue.
<a href="https://habit-hooks.com/" rel="nofollow">https://habit-hooks.com/
I'm using it to foster IOSP (integration operation segregation principle) for example.
fxtentacle · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
DelightOne · · focus · HN ↗
Or is that the secret sauce no one wants to share, the edge people see themselves having.
fxtentacle · · focus · HN ↗
But some of the resellers have freebies, like:
<a href="https://app.primeintellect.ai/dashboard/environments?ex_sort=by_sections" rel="nofollow">https://app.primeintellect.ai/dashboard/environments?ex_sort...
<a href="https://github.com/sierra-research/tau2-bench" rel="nofollow">https://github.com/sierra-research/tau2-bench
balder1991 · · focus · HN ↗