‹ BackHN Continuity

Thread

Dots: Always-on agents

767 points · 647 comments · alvis

  1. mvkel · · focus · HN ↗
    My biggest frustration with the frontier AI companies isn&#x27;t what they&#x27;re announcing, but that the announced-thing that exists ~6 months later is severely nerfed to reduce compute spend. It doesn&#x27;t resemble the demo in any way. For example, this was what the 4o voice capability sounded like in 2024(!) <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=vgYi3Wr7v_g" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=vgYi3Wr7v_g. What exists today pales in comparison.
    1. sleight42 · · focus · HN ↗
      It&#x27;s always about increasing monthly active users, locking them in, and then enshittifying to squeeze out profit.

      Thanks. I&#x27;ll stick with self-hosting.

      1. mvkel · · focus · HN ↗
        I&#x27;d rather use the best possible model available than permanently relegate my work to an inferior one because it&#x27;s &quot;open&quot;
        1. navigate8310 · · focus · HN ↗
          It&#x27;s not only about being open but predictability and unforeseen rug pulling.
          1. mvkel · · focus · HN ↗
            Would you rather work with a team that is hit or miss, but has moments of brilliance, or predictably, consistently bad?
            1. sleight42 · · focus · HN ↗
              Moments of brilliance yet unpredictable versus reasonably good (lets be fair: 27b models are pretty damn good now) and predictable.

              I&#x27;m an old school engineer. I like predictable boring technology that consistently gets the job done over hotness. A 27b model at high quant is &quot;hot&quot; enough.

              Speaking in terms of teams: I hate working with hotshots. They ruin teams. And they&#x27;re usually inconsistent and bad for morale.

        2. calderwoodra · · focus · HN ↗
          Glad someone is saying the quiet part out loud today
      2. kelseydh · · focus · HN ↗
        How do you afford to do this if you want something resembling the best that&#x27;s out there right now? The hardware needed to run beefy open source models is like $15,000 to $50,000+ for a robust local multi-GPU rig, and even its performance might lag behind.
        1. sleight42 · · focus · HN ↗
          I don&#x27;t do that. I use my 2020 top end gaming PC with its 3090. I just ordered 128GB RAM for it. I&#x27;ll use a single-chat&#x2F;slot runtime like Strata or Freetoken. And then I&#x27;ll be able to run either a blazing fast 8-bit Qwen 3.8 27b or a 4 or 5 bit Qwen 3.8 Flash Next. That&#x27;s good enough for me for development.

          And then I&#x27;ll have my current gaming rig with less RAM but better CPU and GPU run a reasonable smart tool agent for handling Home Assistant Voice Assist. Downside there: when I&#x27;m gaming, no voice assist. That may piss the wife off. We&#x27;ll see.

          At least, this way, I don&#x27;t have to worry about god damn usage limits. I can knock myself out.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.