‹ BackHN Continuity

Thread

I built non-autoregressive decision models with RL a year ago

1363 points · 319 comments · nandakishor_ml

  1. edot · · focus · HN ↗
    I don’t understand Jev or this. I used this since it’s open source (good job btw!) with the following. State: “a 6 sided die rolled a 3”, question (noul): “Is the number odd?”

    Answer: 9% chance, with 91% confidence.

    Heh???

    Ok, even worse. 75% chance a coin landed heads up?

    State: I flipped a coin. Question:

    { "noul_result": { "type": "noul", "instructions": "Did the coin land heads up?" }, "choice_result": { "type": "choice", "instructions": "Determine if the coin landed heads or tails up.", "criteria": { "heads": "the coin landed heads up", "tails": "the coin landed tails up" } } }

    Ran on: <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;spaces&#x2F;convaiinnovations&#x2F;laya-demo" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;spaces&#x2F;convaiinnovations&#x2F;laya-demo

    Result: { &quot;model&quot;: &quot;laya&quot;, &quot;answers&quot;: { &quot;noul_result&quot;: { &quot;type&quot;: &quot;noul&quot;, &quot;noul&quot;: 0.6839, &quot;rl_agent&quot;: { &quot;act_probability&quot;: 1.0 } }, &quot;choice_result&quot;: { &quot;type&quot;: &quot;choice&quot;, &quot;choice&quot;: &quot;heads&quot;, &quot;probabilities&quot;: { &quot;heads&quot;: 0.7407, &quot;tails&quot;: 0.2593 }, &quot;confidence&quot;: 0.1743, &quot;rl_agent&quot;: { &quot;act_probability&quot;: 1.0 } } }, &quot;usage&quot;: { &quot;input_tokens&quot;: 76, &quot;output_tokens&quot;: 0 }, &quot;latency_ms&quot;: 93.8 }

    Trying to be even more good-faith:

    State: &quot;A fair coin was flipped once. The result was not observed. No other information about the outcome is available.&quot;

    Questions: { &quot;noul_result&quot;: { &quot;type&quot;: &quot;noul&quot;, &quot;instructions&quot;: &quot;Given only the supplied state, what is the probability that the coin landed heads up?&quot; }, &quot;choice_result&quot;: { &quot;type&quot;: &quot;choice&quot;, &quot;instructions&quot;: &quot;Given only the supplied state, determine which outcome occurred.&quot;, &quot;criteria&quot;: { &quot;heads&quot;: &quot;the coin landed heads up&quot;, &quot;tails&quot;: &quot;the coin landed tails up&quot; } } }

    Result:

    { &quot;model&quot;: &quot;laya&quot;, &quot;answers&quot;: { &quot;noul_result&quot;: { &quot;type&quot;: &quot;noul&quot;, &quot;noul&quot;: 0.1265, &quot;rl_agent&quot;: { &quot;act_probability&quot;: 1.0 } }, &quot;choice_result&quot;: { &quot;type&quot;: &quot;choice&quot;, &quot;choice&quot;: &quot;tails&quot;, &quot;probabilities&quot;: { &quot;heads&quot;: 0.2522, &quot;tails&quot;: 0.7478 }, &quot;confidence&quot;: 0.1853, &quot;rl_agent&quot;: { &quot;act_probability&quot;: 1.0 } } }, &quot;usage&quot;: { &quot;input_tokens&quot;: 123, &quot;output_tokens&quot;: 0 }, &quot;latency_ms&quot;: 154.5 }

    1. bensyverson · · focus · HN ↗
      This is not a good faith test of the system.
      1. edot · · focus · HN ↗
        But it&#x27;s hallucination-free, isn&#x27;t it?
        1. usagisushi · · focus · HN ↗
          yeah, technically. (&#x2F;s)

              python3 - &lt;&lt;&#x27;EOF&#x27;
              import json, urllib.request
              body = json.dumps({
                  &quot;state&quot;: &quot;The car wash is only 100 meters away from my house.&quot;,
                  &quot;model&quot;: &quot;jev-1.13-free&quot;,
                  &quot;questions&quot;: {&quot;q&quot;: {&quot;type&quot;: &quot;choice&quot;,
                      &quot;instructions&quot;: &quot;Should I drive or walk to the car wash?&quot;,
                      &quot;criteria&quot;: {&quot;drive a car&quot;: None, &quot;walk&quot;: None}}}
              }).encode()
              req = urllib.request.Request(&quot;https:&#x2F;&#x2F;opencode.ai&#x2F;zen&#x2F;v1&#x2F;systemone&quot;, data=body,
                  headers={&quot;Content-Type&quot;: &quot;application&#x2F;json&quot;, &quot;User-Agent&quot;: &quot;opencode&#x2F;1.18.31&quot;})
              with urllib.request.urlopen(req, timeout=60) as r:
                  print(json.dumps(json.load(r)[&quot;answers&quot;][&quot;q&quot;], indent=2))
              EOF
              {
                &quot;type&quot;: &quot;choice&quot;,
                &quot;choice&quot;: &quot;walk&quot;,
                &quot;confidence&quot;: 0.66,
                &quot;probabilities&quot;: {
                  &quot;walk&quot;: 0.83,
                  &quot;drive a car&quot;: 0.17
                }
              }
        2. bensyverson · · focus · HN ↗
          I guess we’ve just reached the point where everyone has to state the obvious, and common sense is extremely uncommon.

          So here goes: you should not use an AI model to validate a claim which is trivial to calculate deterministically. That is (obviously?) not what a model like Jev is for, thus it is not a good test of Jev.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.