‹ BackHN Continuity

Thread

Kev: Tiny Jev-like family of decision models built on top of Qwen3.5

462 points · 211 comments · tosh

  1. nico · · focus · HN ↗
    If you only need classification, and you can provide some training data, you can ask Codex/Claude to build an embeddings + logistic classifier model for you

    For emails, I get 95% accuracy with this method, with only 50-100 examples for training

    Training the model takes less than 5 minutes on a CPU

    The resulting model is <1MB, and inference is sub 100ms

    Some other cool things about this approach:

    * the model doesn’t train on some “ideal” or general classification, instead it learns your preferences

    * the model runs on pretty much any mobile device and can be retrained online on the device

    * privacy, the whole training and inference is 100% local, no data goes anywhere (except whatever you feed codex/claude while building the model)

    Note: to do a more general test, I made a classifier for the Banking77 dataset. The model is <10MB, trains in <30s on CPU and gets 94.5% accuracy, which puts it in the top 5?models by accuracy for that set (the best one is at 94.86%, but it’s 350MB in size and takes hours to train on a GPU).

    1. 0x457 · · focus · HN ↗
      I built a whole thing that collects data, trains classifiers, exports models and dataset just for that. Claude writes me a terraform file that contains shape of the classifier and dataset. For images it can create datasets based of another dataset (crop this region from images that have these labels).

      Originally it was so I can label data to fine-tune a VLM, but now a few tiny classifiers that run in milliseconds on cpu.

      Now its collecting data to make a domain specific BERT and do what Jev does.

      1. nico · · focus · HN ↗
        Very cool. What kinda of classifications are you running? How big are the models/training sets?

        Also curious about if you plan on doing some sort of routing for the requests. Like detecting the type of task to decide which model to route the request to

        1. 0x457 · · focus · HN ↗
          Some classifiers are tiny - like 2k params, maybe even less.

          This whole thing started because I wanted something to help me play Dune Imperium. Even relatively large models with vision encoders couldn't reliably extract the full state of the board. Now that I have ~2k labeled screenshots, I want to train heads on top of SigLIP2 to extract all of that data in one go.

          That's how it started. Now the thing supports multiple kinds of datasets:

            Images - currently the Dune Imperium and Bolatro screenshots, with SigLIP2 heads being the next step.
          
            STT - my self-hosted Linux dictation tool feeds this dataset. I run Nemotron ASR tuned for my voice.
          
            TTS - for Piper TTS, trained to speak like SHODAN. Trained from data generated by Qwen3-tts + original video games files.
          
            Text pairs - for a 1.2B model that converts normal text into "what would SHODAN say?"
          
          
            FastApply - a Qwen3.5-4B LoRA adapter for doing fast edits.
          
            Chat threads - all agent/chat threads get saved too, so eventually I can turn the useful ones into a dataset and train a LoRA for a really good Rust-specialized version of Qwen3.8-27B.
          
            Tool calls (extracted from chat threads) - this is where I want something Jev-like, mainly to add an auto-approval mode to my agent harness.
          
          A model router isn't planned because I'm trying to gear everything toward self-hosting, and there just isn't that much to route between. I’ll probably build something Jev-like for smart-home control, though.

          The FastApply dataset is already ~20k entries, with the majority of outputs being 8k–16k tokens. The STT dataset is roughly 30 hours and growing.

          Basically, the whole thing has turned into a Collect -> Distill -> Train pipeline for whatever I happen to need.

          1. nico · · focus · HN ↗
            > Basically, the whole thing has turned into a Collect -> Distill -> Train pipeline for whatever I happen to need

            Amazing, thank you for sharing your setup. Very cool applications

          2. exographicskip · · focus · HN ↗
            > TTS - for Piper TTS, trained to speak like SHODAN

            I didn't know I wanted this! Can just imagine asking how to make a Trinidad Sour and being called an insect lmao

            1. 0x457 · · focus · HN ↗
              It still requires some post-processing to add glitches and metal sound, final result is like this this: <a href="https:&#x2F;&#x2F;drive.google.com&#x2F;file&#x2F;d&#x2F;1CJkK9nHQNBPMr6sIxVmlq8gaxDGwaB0b&#x2F;view?usp=sharing" rel="nofollow">https:&#x2F;&#x2F;drive.google.com&#x2F;file&#x2F;d&#x2F;1CJkK9nHQNBPMr6sIxVmlq8gaxDG...

              I had a version that on validation dataset nearly 1:1 across entire spectrum, but I ended up linking this revision more.

              Also had qwen3-tts version, but no matter what I do it sounded more like brat Cortana. The way piper-tts work under the hood make it sound more machine-like.

              1. andai · · focus · HN ↗
                Nice. I haven&#x27;t played System Shock but this reminded me of the Adjutant from StarCraft.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.