I built non-autoregressive decision models with RL a year ago
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
I built non-autoregressive decision models with RL a year ago
Unofficial Hacker News client; not affiliated with Y Combinator.
woah · · focus · HN ↗
> Zero-shot vs. Fine-tuning: Out-of-the-box base models score ~0.35 on the typed-decisions benchmark (near random). The 0.766 score is achieved by fine-tuning on the benchmark's train split. Treat Laya as a fast foundation model to specialize, not as an omniscient zero-shot oracle.