‹ BackHN Continuity

Thread

Pi Durable

506 points · 72 comments · paulsmith

  1. zmmmmm · · focus · HN ↗
    It's an interesting concept. This is half way to replicating pieces of Gastown. I like the idea, but I'm disappointed these tools still fail to address sandboxing as a first class citizen. I want to be able to declaratively set rules for what sandboxes agents execute in and mark context as tainted when untrusted etc. So far I still don't see any of these harnesses properly addressing this space. I'd be interested in knowing if it can be done through the extensibility of Pi, but since it operates directly on the trust layer, it feels like the type of thing that really needs native support.
    1. jlkuester7 · · focus · HN ↗
      Not familiar with the details of Pi Durable, but I have tinkered a bit with different sandboxing strategies for Pi. IMHO it would be hard to trust a sandboxing layer built into a harness that is so focused on being fully pluggable/moddable/self-improvable.

      When I am using Pi to write extensions for Pi, I feel better running Pi wrapped in a separate os-level sandbox. I guess Pi could do it all, but I am content with how it is.

      1. LeBit · · focus · HN ↗
        After looking at so many options, that is also my take.

        These should be decoupled.

        Maybe I need nono in one context and smolvm in another or both.

        I would not want to trust the harness to self policy.

      2. dbmikus · · focus · HN ↗
        I agree in using a separate OS-level sandbox or a VM. Better to have the option for modularity.

        However, for ease of use, it is nice for harnesses to by default run with sane and safe sandboxing setup. Then give the option to disable them.

      3. zmmmmm · · focus · HN ↗
        if you only have one level of trust then running the harness itself in a sandbox and leaving it at that is fine. This works for coding. For more complex enterprise style scenarios it stops working. Say you have an agent reading emails for you to action high priority ones. You have to assume it is going to get prompt injected constantly. But you want to have an escalation pathway for a high priority email, so somewhere you need a tool that can modify state in a database. You can't give that trust to the email reading one. So you need a higher level agent that can spin up a low trust sub-agent, get an output from it, and then feed the sanitised output into a different agent that has rights to update the database. This is obviously simplified / toy scenario, but it just illustrates that there are different trust levels, and different agents need to be authorised to do different things.
        1. elesiuta · · focus · HN ↗
          I'm currently working on this [1] and can almost support this exact workflow. However the multiple agent orchestration is a sequential state machine, and other than network which can be set for tool states, filesystem access is still set only for the entire state machine.

          [1] <a href="https:&#x2F;&#x2F;github.com&#x2F;agent6-dev&#x2F;agent6" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;agent6-dev&#x2F;agent6

    2. antonok · · focus · HN ↗
      Earendil&#x27;s own Gondolin tool is the best sandboxing model I&#x27;ve found so far. It just executes the toolcalls in a minimal ephemeral VM, unlike most others which run the whole harness inside the sandbox. It&#x27;s a bit rough around the edges (doesn&#x27;t play well with other plugins and doesn&#x27;t work under Bun), but it&#x27;s great if you&#x27;re willing to put in some effort to tweak your setup. Much more comforting to fire off long-running parallel tasks when you know the blast radius is fully contained lol.
      1. patates · · focus · HN ↗
        I&#x27;m not trying to be defeatist but with these models, is there even a real way to contain the blast radius? I also run things sandboxed but it feels like it taking over the whole computer is at the distance of just one probability calculation going awry.
    3. NitpickLawyer · · focus · HN ↗
      &gt; fail to address sandboxing as a first class citizen.

      Isn&#x27;t it better if the tool is sandbox agnostic and you as the developer &#x2F; integrator choose what&#x27;s best for your use case? There are several levels of sandbxing, with many degrees of &quot;freedom&quot;, so it would be really hard&#x2F;confusing&#x2F;overly-complex to build something ootb that suits everyone, no?

    4. whazor · · focus · HN ↗
      The benefit of Pi in my eyes is that the TypeScript interfaces make it easy to build your own sandbox.
    5. badlogic · · focus · HN ↗
      I do not see any resemblence to Gastown at all? Pi Durable is a library for writing durable agents. You can plug in any sandbox solution you like (aka execution environment in Pi Durable speak).

      Within a session, you can give each conversation its own sandbox, based on your application&#x27;s needs and policies.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.