‹ BackHN Continuity

Thread

Turning GLM-5.3-Flash into a Jev-like decision model

138 points · 59 comments · flxflx

  1. walrus01 · · focus · HN ↗
    You can turn any sufficiently smart LLM into yes/no decision model or equivalent. I already have an existing workflow with a two paragraph detailed prompt, that sends pages of stuff to an LLM and asks it to return only 7 JSON objects. Several of those objects are binary "yes or no" choices of like, whether the content contains certain things.

    You can even do it with small not particularly hard to host local LLMs like a variant of Qwen 3.6 35B A3B or 3.8 27B.

    1. opiotrek · · focus · HN ↗
      But does it always 100% of the time sticks to the schema? We have a prompt that is explicitly instructed to return a single html tag with the response inside it and it sometimes hallucinates
      1. walrus01 · · focus · HN ↗
        yes, it does, with appropriate tuning/testing of the llm's parameters (temperature, top p, top k, using the right model, and the right prompt). You have to give it a rigid and very specific prompt to only answer in the JSON form. Also test it with LLMs that will handle being given a very low temperature to be very 'literal'. You're not asking for creative writing.

        I should also add that the source comes from one of about 400 possible places and in a variety of messed up formats, it's the raw feed from a news scraper...

        1. erkl · · focus · HN ↗
          I think you're operating with a different definition of 100%.
          1. walrus01 · · focus · HN ↗
            It hasn't failed once since the single hour I spent writing the prompt and tuning it? Close enough to 100% for my purposes.
            1. krial · · focus · HN ↗
              Yes, but Jev & co literally cannot fail that way. It's 100%. Not the same thing.
              1. walrus01 · · focus · HN ↗
                Cannot fail? It's so smart it literally can't ever parse incoming content incorrectly and provide the wrong answer to a question like "Is at least 25% of the text content in this document written in German? (yes/no)"
                1. david_draco · · focus · HN ↗
                  You are talking about the content of the response, while krial is talking about deviations in the structure of the response. The point is the number of possible states an output can take, so the next step can ingest it and it can be part of a robust pipeline.
      2. gf000 · · focus · HN ↗
        I'm fairly sure it is 100%, and has been available for "ages". That's how every tool call and whatnot works:

        <a href="https:&#x2F;&#x2F;developers.openai.com&#x2F;api&#x2F;docs&#x2F;guides&#x2F;structured-outputs?api-mode=responses" rel="nofollow">https:&#x2F;&#x2F;developers.openai.com&#x2F;api&#x2F;docs&#x2F;guides&#x2F;structured-out...

        1. [deleted] · · focus · HN ↗

          [deleted]

        2. everforward · · focus · HN ↗
          I think the truth is sort of halfway between yall.

          Hallucinations in tool call results _do still exist_, but basically everyone asks for a JSON schema for the tool call and uses that to validate and re-prompt the LLM until it emits something with a valid schema.

          That all goes out the window when a string field has a “hidden” schema in that only particular strings are valid, but that restriction isn’t in the JSON schema. I have had failures when I want a field to be specifically formatted Markdown or something.

          We’ll probably see something that handles this better in the future like jsonnet or Cue or dhall that has some execution capabilities so you can write a custom validator beyond what JSON schema supports.

          1. gf000 · · focus · HN ↗
            &gt; and re-prompt the LLM until it emits something with a valid schema.

            Well, that&#x27;s essential what happens, but on the &quot;backend&quot; side still, so it&#x27;s significantly more efficient. Also, model providers can play with e.g. how likely are they to accept a valid&#x2F;invalid character, so for example the first token may only be [, { or &quot;.

            As for what&#x27;s not in the schema, that&#x27;s an orthogonal question.

          2. hbrn · · focus · HN ↗
            &gt; basically everyone asks for a JSON schema for the tool call and uses that to validate and re-prompt the LLM until it emits something with a valid schema.

            Basically no one is doing that since 2024, have you been living under the rock? Read about constrained decoding.

            &gt; I have had failures when I want a field to be specifically formatted Markdown or something.

            And Jev has an advantage here because it can’t generate anything?

      3. StevenWaterman · · focus · HN ↗
        Yes, you can use constrained decoding
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.