‹ BackHN Continuity

Thread

Dots: Always-on agents

767 points · 647 comments · alvis

  1. jjcm · · focus · HN ↗
    There's a lot of negativity in here for Dots. I've been a pretty heavy user of Grok Bot, and here are a few thoughts a long the positive line.

    1. Collaboration between always-on agents is a really, really powerful thing. It allows for domain-specific expertise that doesn't overload the context window, while still allowing for access to knowledge if they need it.

    2. Domain-specific always on agents creates a good barrier of trust. One of the things I dislike about Claude is sometimes it's memory is all-encompassing. It's weird that it brings up things about my personal life when I'm talking about something related to my business. I've never had that happen with Grok Bot bots because I have one for my biz admin and one for my personal admin. They don't intertwine, which is quite nice.

    3. Combined with cloud agents / cloud builds, things become really powerful for development. It was the first time that I felt there was a solution to the git worktrees / multiple streams at once issue. Each bot has its own computer and can spin up additional cloud agents. It comes at the cost of end to end speed - doing something via a grok bot often takes an hour end to end, whereas with a synchronous local prompt it'll take like 10min. The difference is I have to babysit one whereas the other "just works".

    On the flip side, since using Grok Bots my inference spend has 2-3x'd. It's worth knowing that tradeoff. Nonetheless I think Luna is a fantastic driver for these, and OAI has very good pricing overall. I'd give these a shot - I think a lot of people would be surprised how helpful they are.

    1. mike_hearn · · focus · HN ↗
      Yes. I built my own version of this for my side business about six/seven months ago and it's been great! I have two "AI employees" now and if I were actually focused on this business full time instead of part time, I'd create more.

      Both are just Codexes running in a permanently rolling session in dedicated UNIX user accounts. They're wired up to Maildir so receiving a mail activates Codex and makes it read the new message, there are autonomy wakeup timers, they have accounts in my bug tracker and CI systems. They're currently useful for:

      • Triaging and working on customer support tickets. Sometimes I wake up and the fix/response for a ticket filed by a customer is already there waiting for my approval. Recently I started letting them directly interact with customers in specific scenarios.

      • Triaging the bug backlog. One of them decided to spend its "free time" finding old bugs that were fixed without being properly closed, or are dupes, so it's cleaning up detritus in the tracker.

      • They obviously do all the coding and debugging by just assigning tickets.

      • They keep an eye on a "pet" server the company has, and have proven able to fix it in the past when it ran out of disk space.

      • They handle non-business projects I have for them.

      • They help out with the release processes.

      The dedicated home dir is very useful and they use it all the time as part of coding and investigating tricky issues.

      My setup relies heavily on email, as everything bottoms out in email anyway. Watching them mail each other out of the blue to coordinate stuff is pretty cool.

      1. andai · · focus · HN ↗
        Nice. I also ended up with a Unix user for my agents! (I was looking into Docker etc and realized the only thing I needed was "it doesn't blow up my files", i.e. a linux user).

        I only have one though. What do you have the separate employees for?

        For free time, do you send it mail with cron?

        1. mike_hearn · · focus · HN ↗
          It's to avoid overloading them with disparate tasks and things to keep track of. They have a todo board to help them keep track of things that need doing but there are limits to how far you can push that.

          Another reason: parallelism. The approach of using a single rolling continuously compacting context window is simple and OpenAI are good at compaction, so it works really well. But it means the agent can only do one thing at once. If I send it an email and it decides to spend an hour working on it, then it won't pay attention to any followup emails until after it's done. So having >1 enables more parallelism.

          That said, I don't feel a need for more than two and honestly even that is kind of overkill for the sake of it. For 95% of the time I've been doing this, one was sufficient.

          For free time there are systemd timers that wake it up on a schedule and it uses POSIX locks to mutually exclude runs from different wakeup sources. TODO board items can be either foreground or background; when there's an item with foreground priority the timers wake Codex up a lot more frequently than if there are only background items.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.