‹ BackHN Continuity

Thread

Turning GLM-5.3-Flash into a Jev-like decision model

138 points · 59 comments · flxflx

  1. walrus01 · · focus · HN ↗
    You can turn any sufficiently smart LLM into yes/no decision model or equivalent. I already have an existing workflow with a two paragraph detailed prompt, that sends pages of stuff to an LLM and asks it to return only 7 JSON objects. Several of those objects are binary "yes or no" choices of like, whether the content contains certain things.

    You can even do it with small not particularly hard to host local LLMs like a variant of Qwen 3.6 35B A3B or 3.8 27B.

    1. opiotrek · · focus · HN ↗
      But does it always 100% of the time sticks to the schema? We have a prompt that is explicitly instructed to return a single html tag with the response inside it and it sometimes hallucinates
      1. gf000 · · focus · HN ↗
        I'm fairly sure it is 100%, and has been available for "ages". That's how every tool call and whatnot works:

        <a href="https:&#x2F;&#x2F;developers.openai.com&#x2F;api&#x2F;docs&#x2F;guides&#x2F;structured-outputs?api-mode=responses" rel="nofollow">https:&#x2F;&#x2F;developers.openai.com&#x2F;api&#x2F;docs&#x2F;guides&#x2F;structured-out...

        1. everforward · · focus · HN ↗
          I think the truth is sort of halfway between yall.

          Hallucinations in tool call results _do still exist_, but basically everyone asks for a JSON schema for the tool call and uses that to validate and re-prompt the LLM until it emits something with a valid schema.

          That all goes out the window when a string field has a “hidden” schema in that only particular strings are valid, but that restriction isn’t in the JSON schema. I have had failures when I want a field to be specifically formatted Markdown or something.

          We’ll probably see something that handles this better in the future like jsonnet or Cue or dhall that has some execution capabilities so you can write a custom validator beyond what JSON schema supports.

          1. gf000 · · focus · HN ↗
            &gt; and re-prompt the LLM until it emits something with a valid schema.

            Well, that&#x27;s essential what happens, but on the &quot;backend&quot; side still, so it&#x27;s significantly more efficient. Also, model providers can play with e.g. how likely are they to accept a valid&#x2F;invalid character, so for example the first token may only be [, { or &quot;.

            As for what&#x27;s not in the schema, that&#x27;s an orthogonal question.

          2. hbrn · · focus · HN ↗
            &gt; basically everyone asks for a JSON schema for the tool call and uses that to validate and re-prompt the LLM until it emits something with a valid schema.

            Basically no one is doing that since 2024, have you been living under the rock? Read about constrained decoding.

            &gt; I have had failures when I want a field to be specifically formatted Markdown or something.

            And Jev has an advantage here because it can’t generate anything?

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.