‹ BackHN Continuity

Thread

OpenAI is well positioned to fast-follow Jev

328 points · 233 comments · JohnBerryman

  1. tolugenius · · focus · HN ↗
    I'm not exactly following through with the claim, can someone explain how the built-in classification would not necessitate more tokens used, or be much different from turning on reasoning? Not that I don't see the difference, I just doing see how OpenAI would do it well.
    1. mnicky · · focus · HN ↗
      AFAIK Jev is nothing special technically so it's easy to embed it as an another tool for the LLM? For many batch tasks it can still be quite a token saver I think.

      Or they can even offer it as a standalone API if deemed worth it.

      1. HarHarVeryFunny · · focus · HN ↗
        Jev seems to have three benefits:

        1) It's very cheap and fast - you provide one input and many potential classifications, and the compute to ingest the input is shared.

        2) It generates structured output natively - guaranteed to be correct

        3) It's output probabilities are calibrated to actually mean something

        OpenAI, or anyone else, could certainly replicate it - there are already articles guessing how Jev achieves its "parallel" classifications, but it seems the AI companies need to decide are they in the business of providing intelligence/tokens, or are they in the application business trying to compete with all their customers (not that Jev uses OpenAI).

        1. verdverm · · focus · HN ↗
          (3) seems to be the hard one, you have to have training data with accurate probabilities, maybe, but perhaps not since people are primed to trust
          1. danielmarkbruce · · focus · HN ↗
            No, you don't. You do RLCR, similar to that proposed here:

            <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2507.16806" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2507.16806

            1. verdverm · · focus · HN ↗
              yes, and... pretty much everything in the Ai field comes back to &quot;data makes more difference&quot;
              1. danielmarkbruce · · focus · HN ↗
                Sure, and most days it doesn&#x27;t rain.
                1. verdverm · · focus · HN ↗
                  depends on where you live, an important feature for data points about weather pattern probabilities

                  the underlying data set needs to be representative

                  1. danielmarkbruce · · focus · HN ↗
                    RLVR and RLCR really don&#x27;t need a whole bunch of special data.
                    1. verdverm · · focus · HN ↗
                      the algorithms technically, sure, however the outcomes definitely depend on data quality and coverage like any other training method, this is well known
                      1. danielmarkbruce · · focus · HN ↗
                        I don&#x27;t think you&#x27;ve ever done either of these training steps. You are just handwaving.
                        1. verdverm · · focus · HN ↗
                          you know what they say about making assumptions, yea?

                          and then you are going to ignore all the research and results that clearly show otherwise? why?

                          what might we infer about the importance of data from a learning algorithm like decision trees?

                          1. danielmarkbruce · · focus · HN ↗
                            Read the paper. They train RLCR on existing big math problems. They subtract a brier score penalty from the correctness reward. No new confidence labels are needed.

                            Existing datasets, different reward function.

                            1. verdverm · · focus · HN ↗
                              &gt; Read the paper.

                              I did, in the first days Jev came out, when people were bringing it up. Another assumption. Please review the HN commenting guidelines, the one which starts with &quot;Please don&#x27;t comment on whether someone read an article.&quot; is relevant here.

                              Nothing in that paper changes that ML algorithms are dependent on the training data. We can step back from Jev and algos to consider Bayes Theorem. If your sample is not representative of the population, your resulting statistics will be off. The same is true here. If the data you train a model like Jev with is not representative, the probabilities and confidences it outputs will not be representative.

                              What makes Jev interesting is that it works well out of the box across domains. What people who are well known in the field believe is that this is the result of Typesafe having a really good training data set. People are saying similar of MiMo-2.6 today.

                              1. danielmarkbruce · · focus · HN ↗
                                &quot;Did you read the article&quot; doesn&#x27;t apply to a link someone put in a comment. If you are going to be a hall monitor, at least do it properly. You are just acting in bad faith at this point.
                                1. verdverm · · focus · HN ↗
                                  You are not engaging with actual points, instead attacking a person based on your bad assumptions and projections.

                                  We both know who is

                                  &gt; just acting in bad faith at this point.

                                  1. danielmarkbruce · · focus · HN ↗
                                    The relevant data is the reasoning trace. Doesn&#x27;t need user data. You can learn from people&#x27;s detailed reasoning steps how confident they are, even outside your domain.

                                    Take RL 101. This is a common pattern.

                                    1. verdverm · · focus · HN ↗
                                      We were talking about Jev and probability, now you&#x27;re changing the problem, a rhetorical trick some people try to employ.

                                      Another that uses dice rolling, coin flips, and an inventory level example to drive home the point that Jev&#x27;s output are not real probabilities for outcomes.

                                      <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49830385">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49830385

                                      &gt; Take RL 101

                                      I taught it (ML course; a day on RL, at a university), you should really stop making assumptions friend. Data quality and coverage matters in learning algorithms.

                                      Here&#x27;s one of the books used in that course <a href="https:&#x2F;&#x2F;amlbook.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;amlbook.com&#x2F;

                                      Thinking blocks are not a place you can derive real confidence scores in LLMs

                                      1. danielmarkbruce · · focus · HN ↗
                                        My initial comment and every one following is about RLCR and that paper. You don&#x27;t appear to grasp the basics of that paper, it&#x27;s reward function or how the optimizer is updating weights.

                                        You are out of your depth and grasping at straws.

                                        1. verdverm · · focus · HN ↗
                                          &gt; You are out of your depth and grasping at straws.

                                          Do you have any credentials or evidence that others can use to determine if this statement is not more accurately describing the author who wrote it?

                                          Perhaps a PhD in ML, research output like published papers, or teaching&#x2F;professional experience - all things I have

                                          We could debate the merits of the paper contents, but I suspect you have intentionally moved on to personal attacks. Regardless, nothing you have said (nor can be found in this paper) has been a counter argument that learning algorithms are sensitive to training data, where the measured output difference is used by the optimization algorithm when updating the parameters. Garbage in, garbage out is a saying for a reason. No algorithm fixes non-representative data.

                                          1. danielmarkbruce · · focus · HN ↗
                                            &gt;you have to have training data with accurate probabilities

                                            This was your claim. If you can&#x27;t read and understand that paper in relation to your claim, you are out of your depth. You haven&#x27;t made a single claim relevant to that paper - just hand wavy comments about data.

                                            1. verdverm · · focus · HN ↗
                                              I clarified multiple times that I meant &quot;accurate, representative data&quot; to &quot;derive accurate probabilities&quot;

                                              you are still employing underhanded techniques in an attempt &quot;win an internet debate&quot; (my impression)

                                              try being more accommodating and flexible over repeating the same lame things

                                              it&#x27;s not hard to say, &quot;ah I see what you were trying to say...&quot; and move towards a more constructive conversation

                                              RLCR &#x2F; Jev et al. can only give as accurate predictions and probabilities as the underlying data they are trained on represents. Biased data results in biased probabilities, no algorithm fixes this. Can we agree on this point?

                                              1. danielmarkbruce · · focus · HN ↗
                                                They post train on big math and improve calibration across five different non math benchmarks so the new claim &quot;can only give as accurate predictions and probabilities as the underlying data they are trained on represents&quot; is also off base. RLCR generates its own calibration examples from ordinary questions and answer keys. RL usually isn&#x27;t trying to represent a data set, it&#x27;s closer to search.
                                                1. verdverm · · focus · HN ↗
                                                  I&#x27;m not expecting you to human RL here, but I do hope next time you remember there is another imperfect human behind the screen and are less too online in future interactions

                                                  <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=c1Fv1uKTd-w" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=c1Fv1uKTd-w

                                                  oh-seven

                                                  1. danielmarkbruce · · focus · HN ↗
                                                    A lot of pretzeling to avoid understanding one paper.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.