‹ BackHN Continuity

Thread

The OpenAI Decisions API needs a confidence you can trust

33 points · 15 comments · AnthusAI

  1. rileymat2 · · focus · HN ↗
    Alan is young, round, and kind, but that doesn't mean he isn't also rough and cold at times, as well. … Young round people who are green are usually blue. … Kind people with rough skin are usually red because it's wind burn. If someone shows that they are red, then they are also showing that they are green. …

    Statement: Alan is not blue.

    A log-probability of −0.00182 is a probability of 99.82%. We asked for five alternatives and got none: Luna put essentially nothing on "true" or "false". And it's wrong. Alan is kind with rough skin, so he's red; red means green; young, round and green means blue. "Alan is not blue" is false, three steps in.

    —————-

    Can someone explain this I got unknown as well. The problem statement includes the word “usually” a few times.

    1. perching_aix · · focus · HN ↗
      It took me 15 minutes to decipher this and arrive at an interpretation that both isn't incoherent nonsense, and actually answers the question correctly.

      If this trips up decision models that operate on the scale of tens to hundreds of milliseconds, maybe that's okay? This is super contrived, like all riddles are.

      1. rileymat2 · · focus · HN ↗
        I could not get to the expected result, maybe I Was thinking about absolute truths, but the statement you are expected to answer is an absolute, but built on "usually" factors.

        The explanation says : "Alan is kind with rough skin, so he's red". But they ignored "usually".

        So for instance I could say country music fans are usually white. Darius Rucker is a country fan.

        Statement: Darius Rucker is white.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.