you don't need to post-train anything for this.
just get an LLM to think and then force it to output a specific json with prefill post-think.
make sure to include good conditioning text in the prompt with examples of exactly what the output should be like. you don't want dissonance in the probabilities on the prefill.
A classifier is a subset of generative text, so I think responses like this miss the point. Jev is cheap enough and fast enough to sprinkle across your app in ways that LLM would be infuriatingly laggy and unnecessarily expensive, and it's never going to be injected to provide a sorting algorithm in python.
The point isn't that new type of problem has been unlocked, rather a new approach that can unlock new use cases.
this is exactly why Jev doesn't have thinking.
when you want a machine to reason about the prompt and generate a structured output not using an actual LLM makes no sense. I have been doing it since the first chain of thought open models became available.
perhaps there may be a way to get a Jev-type model to think for a very specific number of steps to gain control over its latency, if so that would be the next step. truncating LLM thinking like this does not work well, and its thinking isn't efficient anyway.
teravor · · focus · HN ↗
just get an LLM to think and then force it to output a specific json with prefill post-think.
make sure to include good conditioning text in the prompt with examples of exactly what the output should be like. you don't want dissonance in the probabilities on the prefill.
betenoire · · focus · HN ↗
The point isn't that new type of problem has been unlocked, rather a new approach that can unlock new use cases.
teravor · · focus · HN ↗
when you want a machine to reason about the prompt and generate a structured output not using an actual LLM makes no sense. I have been doing it since the first chain of thought open models became available.
perhaps there may be a way to get a Jev-type model to think for a very specific number of steps to gain control over its latency, if so that would be the next step. truncating LLM thinking like this does not work well, and its thinking isn't efficient anyway.