‹ BackHN Continuity

Thread

Turning GLM-5.3-Flash into a Jev-like decision model

138 points · 59 comments · flxflx

  1. walrus01 · · focus · HN ↗
    You can turn any sufficiently smart LLM into yes/no decision model or equivalent. I already have an existing workflow with a two paragraph detailed prompt, that sends pages of stuff to an LLM and asks it to return only 7 JSON objects. Several of those objects are binary "yes or no" choices of like, whether the content contains certain things.

    You can even do it with small not particularly hard to host local LLMs like a variant of Qwen 3.6 35B A3B or 3.8 27B.

    1. opiotrek · · focus · HN ↗
      But does it always 100% of the time sticks to the schema? We have a prompt that is explicitly instructed to return a single html tag with the response inside it and it sometimes hallucinates
      1. walrus01 · · focus · HN ↗
        yes, it does, with appropriate tuning/testing of the llm's parameters (temperature, top p, top k, using the right model, and the right prompt). You have to give it a rigid and very specific prompt to only answer in the JSON form. Also test it with LLMs that will handle being given a very low temperature to be very 'literal'. You're not asking for creative writing.

        I should also add that the source comes from one of about 400 possible places and in a variety of messed up formats, it's the raw feed from a news scraper...

        1. erkl · · focus · HN ↗
          I think you're operating with a different definition of 100%.
          1. walrus01 · · focus · HN ↗
            It hasn't failed once since the single hour I spent writing the prompt and tuning it? Close enough to 100% for my purposes.
            1. krial · · focus · HN ↗
              Yes, but Jev & co literally cannot fail that way. It's 100%. Not the same thing.
              1. walrus01 · · focus · HN ↗
                Cannot fail? It's so smart it literally can't ever parse incoming content incorrectly and provide the wrong answer to a question like "Is at least 25% of the text content in this document written in German? (yes/no)"
                1. david_draco · · focus · HN ↗
                  You are talking about the content of the response, while krial is talking about deviations in the structure of the response. The point is the number of possible states an output can take, so the next step can ingest it and it can be part of a robust pipeline.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.