‹ BackHN Continuity

Thread

I built non-autoregressive decision models with RL a year ago

1363 points · 319 comments · nandakishor_ml

  1. srameshc · · focus · HN ↗
    from <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;convaiinnovations&#x2F;laya" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;convaiinnovations&#x2F;laya &gt; The policy reports a distribution; exploration adds zero-mean Gaussian noise to the logits; the reward is a strictly proper scoring rule (log + spherical, plus ranked probability score for ordinal questions). Expected reward is maximised only by reporting honest probabilities.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.