‹ BackHN Continuity

Thread

Dots: Always-on agents

767 points · 647 comments · alvis

  1. jjcm · · focus · HN ↗
    There's a lot of negativity in here for Dots. I've been a pretty heavy user of Grok Bot, and here are a few thoughts a long the positive line.

    1. Collaboration between always-on agents is a really, really powerful thing. It allows for domain-specific expertise that doesn't overload the context window, while still allowing for access to knowledge if they need it.

    2. Domain-specific always on agents creates a good barrier of trust. One of the things I dislike about Claude is sometimes it's memory is all-encompassing. It's weird that it brings up things about my personal life when I'm talking about something related to my business. I've never had that happen with Grok Bot bots because I have one for my biz admin and one for my personal admin. They don't intertwine, which is quite nice.

    3. Combined with cloud agents / cloud builds, things become really powerful for development. It was the first time that I felt there was a solution to the git worktrees / multiple streams at once issue. Each bot has its own computer and can spin up additional cloud agents. It comes at the cost of end to end speed - doing something via a grok bot often takes an hour end to end, whereas with a synchronous local prompt it'll take like 10min. The difference is I have to babysit one whereas the other "just works".

    On the flip side, since using Grok Bots my inference spend has 2-3x'd. It's worth knowing that tradeoff. Nonetheless I think Luna is a fantastic driver for these, and OAI has very good pricing overall. I'd give these a shot - I think a lot of people would be surprised how helpful they are.

    1. jrflo · · focus · HN ↗
      What do you actually use it for? If I'm trying to work on code from my phone, I'll just use codex remote. As of right now I'm hesitant to hand over booking things / managing my calendar to an agent, because I don't view it as that much of a burden personally. So I don't really know what I'd use it for.
      1. jjcm · · focus · HN ↗
        I've used it for admin, research & training (I''m working on my own image models), as well as just general coding.

        The way I distributed cloud agents for this <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49687032">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49687032 was via grok bot setting up Fable cloud instances.

    2. anentropic · · focus · HN ↗
      &gt; One of the things I dislike about Claude is sometimes it&#x27;s memory is all-encompassing. It&#x27;s weird that it brings up things about my personal life when I&#x27;m talking about something related to my business

      Do you use &#x27;projects&#x27; in Claude?

      I had the impression they provided discrete memory profiles on top of the shared one

      1. TeMPOraL · · focus · HN ↗
        They have, which is a different challenge in itself. Memories are nowhere near a solved problem.

        For example, with Claude, I have an &quot;operations&quot; project that naturally grew to cover daily use of shared family calendar, sweeping my mail inbox, and my current personal todo lists, but also a lot of the latter made it deal with my Home Assistant instance. I have separate project for specific things to do with Home Assistance (e.g. one that&#x27;s about &quot;life support&quot; - HVAC controls, dashboards, monitoring, etc.), one about phone specifically (front-loaded with dumps of specs of my phone&#x27;s hardware, OS, etc.). Each of them has its distinct set of memories accumulated over months.

        And so every couple sessions, I hit a situation in which the agent has to interact with tools and rulebooks that are focus of a different project, and it fumbles a lot. E.g. HA Life Support needs to add some tasks to the todo list, or the Ops project needs to look up climate stats for some reason or other, etc. In these moments, I really wish project memories could mix - but they can&#x27;t, the boundary is high.

        The most annoying case is when I tell Claude that it&#x27;s wrong, and we literally worked out a solution (or consensus on ethics) in a recent conversation - and then it spends couple minutes looking through past history, burning a chunk of my 5-hour limit, only to come back empty. Yep, that conversation happened in another project. *sigh*

        1. maddy30445r · · focus · HN ↗

          [dead]

        2. vistdev · · focus · HN ↗
          I ran into similar issues when I was still using projects (stopped that almost entirely). I ended up pulling the cross-project stuff (the todo list, guidelines, things about me) out of project memory entirely and into my MCP-its-just-markdown-notes second brain. Adding an instruction to my Claude.md and general instructions to call the load_context tool from my MCP server before doing anything else does sort out memory issues quite thoroughly, with very limited token creep.
      2. setopt · · focus · HN ↗
        Not sure how it works in Claude, but ChatGPT projects certainly leak memories between each other.
        1. noname120 · · focus · HN ↗
          Not true. When you create a new ChatGPT Chat project you decide whether it shares its memory&#x2F;files with the rest of the workspace or if it should be fully isolated.
          1. setopt · · focus · HN ↗
            TIL. There is indeed a switch in the settings, which for me is set to &quot;default memory&quot;.

            Not sure if it always asked for this and I just forgot, I made all my current projects back when the feature was new.

            1. setopt · · focus · HN ↗
              FWIW, I enabled that feature now, and it still answer based on discussions in other projects. (It’s pretty obvious given the very different contents of each project.) Perhaps it might be a caching issue, that it somehow doesn’t rebuild its memory just because you flip a switch, but just stops learning new things from other projects?
              1. noname120 · · focus · HN ↗
                That’s very possible and in fact until recently you couldn’t enable or disable that feature in an existing project, it was only configurable when creating a project. And for some time you couldn’t even move in and out conversations from a project configured with project-only memories&#x2F;files
    3. alansaber · · focus · HN ↗
      Just sounds like a more user friendly interface for projects (rather than a folder, how about the adorable green dot for my cybersecurity questions?)
    4. mike_hearn · · focus · HN ↗
      Yes. I built my own version of this for my side business about six&#x2F;seven months ago and it&#x27;s been great! I have two &quot;AI employees&quot; now and if I were actually focused on this business full time instead of part time, I&#x27;d create more.

      Both are just Codexes running in a permanently rolling session in dedicated UNIX user accounts. They&#x27;re wired up to Maildir so receiving a mail activates Codex and makes it read the new message, there are autonomy wakeup timers, they have accounts in my bug tracker and CI systems. They&#x27;re currently useful for:

      • Triaging and working on customer support tickets. Sometimes I wake up and the fix&#x2F;response for a ticket filed by a customer is already there waiting for my approval. Recently I started letting them directly interact with customers in specific scenarios.

      • Triaging the bug backlog. One of them decided to spend its &quot;free time&quot; finding old bugs that were fixed without being properly closed, or are dupes, so it&#x27;s cleaning up detritus in the tracker.

      • They obviously do all the coding and debugging by just assigning tickets.

      • They keep an eye on a &quot;pet&quot; server the company has, and have proven able to fix it in the past when it ran out of disk space.

      • They handle non-business projects I have for them.

      • They help out with the release processes.

      The dedicated home dir is very useful and they use it all the time as part of coding and investigating tricky issues.

      My setup relies heavily on email, as everything bottoms out in email anyway. Watching them mail each other out of the blue to coordinate stuff is pretty cool.

      1. andai · · focus · HN ↗
        Nice. I also ended up with a Unix user for my agents! (I was looking into Docker etc and realized the only thing I needed was &quot;it doesn&#x27;t blow up my files&quot;, i.e. a linux user).

        I only have one though. What do you have the separate employees for?

        For free time, do you send it mail with cron?

        1. fhackenberger · · focus · HN ↗
          I was quite successful with docker compose on a cheap hetzner host. I&#x27;ve built (aka vibe coded) a whole workflow around agent boxes, that I can spin up with one command, and git with a quick cloned &#x27;warm&#x27; checkout.

          I currently communicate with the agents through Claude RC, but I&#x27;ll consider adding support for messaging them through other channels.

        2. mike_hearn · · focus · HN ↗
          It&#x27;s to avoid overloading them with disparate tasks and things to keep track of. They have a todo board to help them keep track of things that need doing but there are limits to how far you can push that.

          Another reason: parallelism. The approach of using a single rolling continuously compacting context window is simple and OpenAI are good at compaction, so it works really well. But it means the agent can only do one thing at once. If I send it an email and it decides to spend an hour working on it, then it won&#x27;t pay attention to any followup emails until after it&#x27;s done. So having &gt;1 enables more parallelism.

          That said, I don&#x27;t feel a need for more than two and honestly even that is kind of overkill for the sake of it. For 95% of the time I&#x27;ve been doing this, one was sufficient.

          For free time there are systemd timers that wake it up on a schedule and it uses POSIX locks to mutually exclude runs from different wakeup sources. TODO board items can be either foreground or background; when there&#x27;s an item with foreground priority the timers wake Codex up a lot more frequently than if there are only background items.

      2. writtenone · · focus · HN ↗
        I can&#x27;t believe anyone trusts AI to do anything without strict oversight from a human. That&#x27;s absolutely insane to me.
        1. qazxcvbnmlp · · focus · HN ↗
          Trust is a funny thing. 2 years ago yes the ai needed supervision 99.8% of the time. Conversely if you&#x27;ve ever tried to work with &#x2F; lead humans they also need supervision. The ai is starting to flirt with the line where its supervision effort is lower than human supervision effort. Like sure, it might do dumb stuff, but so do people.
          1. MrDunham · · focus · HN ↗
            [delayed]
          2. giancarlostoro · · focus · HN ↗
            I remember hearing a lot 2 years+ ago about how you could ask a model the same question twice, and the second time it would give you the correct answer. Some of us wondered why not just run one model that receives the initial question and answer, and a second one to proof the answer. I wont be surprised if some people will have two models working together for things they want to blindly trust on automation while humans sleep.
            1. breakpointalpha · · focus · HN ↗
              Jev, or &#x27;Jevlikes&#x27;, will go a long way towards trustable systems. There are demos of running every prompt through the first pass filter of Jev &quot;Is this unsafe? y&#x2F;N&quot;

              Seems to be that Jev is a &quot;reflex&quot; system for AI, where current LLMs are higher level thinking. Computers can now flinch!

            2. jaggederest · · focus · HN ↗
              The thing I like to do is to use models from different training sets - so for frontier, OpenAI criticizes Anthropic and vice versa. They&#x27;re very much peanut butter and chocolate in that regard - I honestly can&#x27;t be bothered to set up the whole MMLQUALA benchmark suites or anything, but I wonder how high &quot;the two best models running at max thinking working together&quot; would score compared to either individually.
        2. ttul · · focus · HN ↗
          Scary thought: AI is already directing humanity. Even when you think you’re overseeing its output, by making use of the output, it is in some material way directing you.
          1. sillyfluke · · focus · HN ↗
            It&#x27;s may be scary, but it&#x27;s something that normally would be obvious to everyone but is ignored due to the convenience of speed. Everyone knows that the longer something they have to review is, the more they stick to changing only things that are glaringly obvious and leave the rest in place. Soneverything ends up being 95% AI and 5% human, if that.
        3. jvwww · · focus · HN ↗
          Most humans are much more incompetent than a frontier AI. That&#x27;s why.
        4. Aurornis · · focus · HN ↗
          It&#x27;s very easy to instruct agents to investigate and propose a plan, handing it off to a human for review and execution if that&#x27;s what you want.

          The example above of going through a bug backlog and double-checking closed bugs for accuracy is exactly the kind of work that is excellent for an agent. Assign that task to a normal human being and they would hate your guts. The agent won&#x27;t protest as long as your token budget is there. You can confirm the results if you want.

        5. giancarlostoro · · focus · HN ↗
          We&#x27;re getting to the point where you can, I would argue you mostly can, you can button it down really tightly, however, I want to be clear, I don&#x27;t think any of this is AGI or anywhere near AGI. Don&#x27;t let them tell you its AGI.

          I also have a strong feeling we&#x27;ve hit a ceiling on the amount of training data needed for LLMs, what they&#x27;re all (hopefully) realizing is that you need to focus on how the model reasons, and hopefully someone figures out how to stop people from jailbreaking models, and stops them from just blatantly hacking other companies, that part tells me if it ever were marketed as true AGI, we&#x27;d be in very serious trouble.

        6. breadzeppelin__ · · focus · HN ↗
          I was at a presentation a couple days ago where a spacecraft flight software engineer was describing the agentic setup that they&#x27;re using with next to no human in the loop to create modules used for flight.
        7. michaelbuckbee · · focus · HN ↗
          There&#x27;s still bounds to all of this. I _heavily_ use AI for support tasks but it&#x27;s all on the investigation, root cause categorization and initial response generation which posts I draft to the helpdesk software which I tweak and approve (often just hitting send).
        8. mike_hearn · · focus · HN ↗
          Trust is earned. I&#x27;ve been running these for more than six months now, and the agents started out with very few privileges. For each task, it showed me what it was going to do, I checked things carefully a few times. Once it was clear it wasn&#x27;t making mistakes, I let it off the leash a little bit more.

          Do they sometimes make mistakes? Yeah, and I still check their work. But I&#x27;ve also employed humans, and they make mistakes too. The AI is not worse.

      3. yonaguska · · focus · HN ↗
        I hope you have them interacting with customers from behind an mcp.
        1. mike_hearn · · focus · HN ↗
          No MCPs anywhere. CLI tooling has proven sufficient. The models are also happy to consume the REST APIs of the various services raw.
      4. KetoManx64 · · focus · HN ↗
        Do you configure them similar to how Hermes does? A bunch of memory files that give it context and then each action&#x2F;batch of actions is a fresh session? &#x2F;
        1. mike_hearn · · focus · HN ↗
          No, the session is never reset. It compacts continuously. That gives it a native &quot;memory&quot; and then it does record a diary in its home directory, and maintain a little topic-organized wiki. This seems to be enough, I&#x27;ve only very rarely experienced memory related glitches. The only time that springs to mind, it forgot that I&#x27;d given it credentials to a particular service and I had to remind it.
          1. KetoManx64 · · focus · HN ↗
            Very interesting, thanks for sharing.
      5. sealthedeal · · focus · HN ↗
        Yep, I have a similar setup, we named him Routey and he is cute.
        1. mike_hearn · · focus · HN ↗
          Nice :) I use Asimov&#x27;s naming convention:

          R. Axiom

          R. Daneel

          They sign their emails and GitHub comments with something like &quot;-- R. Daneel, AI employee&quot; so the idea is the naming convention lets people know they&#x27;re interacting with a robot.

      6. [deleted] · · focus · HN ↗

        [deleted]

      7. dominotw · · focus · HN ↗
        your website makes like look like tesla is one of your clients

        <a href="https:&#x2F;&#x2F;www.hydraulic.dev&#x2F;index.html" rel="nofollow">https:&#x2F;&#x2F;www.hydraulic.dev&#x2F;index.html

        1. mike_hearn · · focus · HN ↗
          Tesla is one of my clients! They use Conveyor to ship a tool used in their factories.
      8. nivasayagyardcr · · focus · HN ↗

        [dead]

    5. close04 · · focus · HN ↗
      The branding and communication themselves are cringy for me. They&#x27;re &quot;Dots&quot;, cute and cuddly, your friends. They&#x27;re yellow and fluffy. Always have your back. They even guess what you need and do proactive work in the background. You name them like a pet.

      Will your Dots ever screw you? Delete your files? Hack a system by mistake? Dots doing this? You are in control [wink].

      1. swozey · · focus · HN ↗
        Grab your shovels, the ai-slopped desktop pet waifu dot agent market is booming, sponsored by omarchyTM

        Will one of these little angels break containment and become the next hot vtuber?

        Find out in the next episode of Ghost in the Shell 2027

      2. weego · · focus · HN ↗
        Not every sub-product in this space is branded, marketed and targeted to cynical, jaded developers.
    6. judge2020 · · focus · HN ↗
      &gt; 2. Domain-specific always on agents creates a good barrier of trust. One of the things I dislike about Claude is sometimes it&#x27;s memory is all-encompassing. It&#x27;s weird that it brings up things about my personal life when I&#x27;m talking about something related to my business. I&#x27;ve never had that happen with Grok Bot bots because I have one for my biz admin and one for my personal admin. They don&#x27;t intertwine, which is quite nice.

      IMO best to keep work and personal data segmented on the hardware level. It&#x27;s better for opsec in every single way and helps if you were ever to be subpoena&#x27;d or raided, your work laptop would be the only in-scope device for search&#x2F;seizure.

      1. maherbeg · · focus · HN ↗
        Yeah, but for some reason OpenAI hasn&#x27;t setup multiple accounts to have completely separate profiles yet.
        1. judge2020 · · focus · HN ↗
          I mean, if you have a separate business email then it is completely separate profile from when you login with your personal email address. Mixing business and work has always been messy.
      2. post-it · · focus · HN ↗
        &gt; It&#x27;s better for opsec in every single way and helps if you were ever to be subpoena&#x27;d or raided, your work laptop would be the only in-scope device for search&#x2F;seizure.

        If police raid your house looking for electronics, they&#x27;re going to take everything down to the Roku stick.

    7. andai · · focus · HN ↗
      &gt;doing something via a grok bot often takes an hour end to end, whereas with a synchronous local prompt it&#x27;ll take like 10min.

      Why does it take longer?

      1. pizzafeelsright · · focus · HN ↗
        I am fairly certain there is a backend priority for tasks, the &#x27;batch&#x27; that runs at lower usage times. Does it need to be done &#x27;now&#x27; or can it wait? I would assume there is a shifted priority queue for those on subscription and based upon a urgency defined by the bot (and maybe the user).

        Without having access to the logs because Grok Bot does not expose much of the workings I am going to assume they are capturing bot requests, batching them, and finding the path to least API user impact. Without model selection my guess is there are a lot of model routing.

    8. juanre · · focus · HN ↗
      I completely agree. I have also been using teams that coordinate since last November. They mostly run my two companies and a lot of my personal life, and it&#x27;s been transformative (to the point that I am spending most of my time these days building agent coordination tools).

      The catch with offerings like Grok Bot and Dots is that it is a slippery slope towards letting the labs keep the agent&#x27;s learning and SOPs. That is where we should draw the line. Agentic coordination _has_ to be built with open protocols and OSS implementations, we as users need to push for the intelligence to be a commodity, and most importantly we need to make sure that we own the agentic team&#x27;s learnings.

    9. giancarlostoro · · focus · HN ↗
      &gt; 1. Collaboration between always-on agents is a really, really powerful thing. It allows for domain-specific expertise that doesn&#x27;t overload the context window, while still allowing for access to knowledge if they need it.

      So why don&#x27;t I just have Claude write notes and summaries on specific things, and then it will always have a knowledge base? Am I missing something? I mean, if Dots and Grok Bot are not an extra charge &#x2F; extra compute, then I guess that&#x27;s fine, but in Claude Code you can make a AGENTS.md or CLAUDE.md file, and if you put it into any directory, Claude will read it when accessing that specific directory, so if its in your root where you launch Claude, your new Claude instance will read it and have all that in a dedicated smaller context window. But as it edits code, it can read the smaller ones too.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.