Can someone explain how these durable agents handle state present on the VMs/Sandboxes ? I get it that the agent state can be recreated from checkpoints/logs, but what about the state present on the runners (i.e. container, VMs, Sandboxes, etc). How are both states kept in sync ?
Like if I have a web-app running on the runner and the agent is navigating the web UI and then the runner (or the agent) crashes. When the agent is recreated back from the checkpoints (or a new runner is launched), it will think it has already navigated to page N, but in reality the browser on the runner might be on page 0.
Think about how you'd do it as a human, that's usually the answer for these things in my experience.
If you've ever worked with any workflow engine, doing it with agents is largely the same. If you were writing some automation that used a browser, how would you handle recovery for any given step of your workflow? It depends on what you're doing, the specifics of the web app you're interfacing with, etc.
Contrived example, but let's say you're sending an email. Load the page, click the button, enter text in the various fields, etc. Since there's no side effect of consequence until you hit send, you could just make sure your failures clean up drafts, and replaying the whole thing is safe.
I've admittedly done very little browser automation like this though, mainly I've just called APIs, created and uploaded files, done db operations, normal dev stuff.
I'd expect RPA platforms to be way ahead on agent automation like this, I haven't kept with any of them though. If they aren't, real missed opportunity for them.
The good version of it is to only allow interacting with the environment declaratively (preferably idempotently as well), and then checkpointing those.
So rather than saying “open chrome, go to this page, click next page 5 times”, it would be something like “chrome is running; url is X; url is X/page/1; url is X/page/2” etc. Ansible, basically.
Other than that, you could just replay bash tool calls. That’s full of holes, though. Anything that relies on “date” will return different stuff, and if you try checkpointing the system time then TLS breaks due to timestamp differences.
If you wanted to go absolutely wild, some hypervisors can checkpoint the memory of a running VM and revert back to a prior version memory and all. I can’t imagine a way to make money off that (you’d be writing gigs of data per checkpoint), but I suppose it’s technically possible.
> If you wanted to go absolutely wild, some hypervisors can checkpoint the memory of a running VM and revert back to a prior version memory and all. I can’t imagine a way to make money off that
You may want to check out what Antithesis is doing in this space - <a href="https://antithesis.com/" rel="nofollow">https://antithesis.com/ Specifically, <a href="https://antithesis.com/docs/resources/deterministic_simulation_testing/" rel="nofollow">https://antithesis.com/docs/resources/deterministic_simulati...
It’s neat but I don’t think fixes these issues because they have to integrate with live systems. You can lie about system time in tests by making everything else have the same time.
That doesn’t apply if I need to hit gmail.com and can’t login because my system time is a week behind and the JWT says it isn’t valid for another week.
Even if you make everything else accept the time, timestamps will be screwed. Like in a fake gmail service that accepts mocked times, what do you use for timestamps on emails?
With the one I’ve made, it’s best effort. We record whether or not tool calls have succeeded and if the agent dies without knowing whether or not the call succeeded, then when we resurrect it we tell it the state is unknown and it can either inspect the resource or try again depending on the side effects.
> Can someone explain how these durable agents handle state present on the VMs/Sandboxes ?
That greatly depends on your agent design. If you give a user a full sandbox then you're going to be in a position where you probably need to snapshot it. But there are plenty of agent designs that are not using full VMs and for those the state story is way easier.
pulkitsh1234 · · focus · HN ↗
Like if I have a web-app running on the runner and the agent is navigating the web UI and then the runner (or the agent) crashes. When the agent is recreated back from the checkpoints (or a new runner is launched), it will think it has already navigated to page N, but in reality the browser on the runner might be on page 0.
arnorhs · · focus · HN ↗
phoghed · · focus · HN ↗
If you've ever worked with any workflow engine, doing it with agents is largely the same. If you were writing some automation that used a browser, how would you handle recovery for any given step of your workflow? It depends on what you're doing, the specifics of the web app you're interfacing with, etc.
Contrived example, but let's say you're sending an email. Load the page, click the button, enter text in the various fields, etc. Since there's no side effect of consequence until you hit send, you could just make sure your failures clean up drafts, and replaying the whole thing is safe.
I've admittedly done very little browser automation like this though, mainly I've just called APIs, created and uploaded files, done db operations, normal dev stuff.
I'd expect RPA platforms to be way ahead on agent automation like this, I haven't kept with any of them though. If they aren't, real missed opportunity for them.
everforward · · focus · HN ↗
So rather than saying “open chrome, go to this page, click next page 5 times”, it would be something like “chrome is running; url is X; url is X/page/1; url is X/page/2” etc. Ansible, basically.
Other than that, you could just replay bash tool calls. That’s full of holes, though. Anything that relies on “date” will return different stuff, and if you try checkpointing the system time then TLS breaks due to timestamp differences.
If you wanted to go absolutely wild, some hypervisors can checkpoint the memory of a running VM and revert back to a prior version memory and all. I can’t imagine a way to make money off that (you’d be writing gigs of data per checkpoint), but I suppose it’s technically possible.
svieira · · focus · HN ↗
You may want to check out what Antithesis is doing in this space - <a href="https://antithesis.com/" rel="nofollow">https://antithesis.com/ Specifically, <a href="https://antithesis.com/docs/resources/deterministic_simulation_testing/" rel="nofollow">https://antithesis.com/docs/resources/deterministic_simulati...
everforward · · focus · HN ↗
That doesn’t apply if I need to hit gmail.com and can’t login because my system time is a week behind and the JWT says it isn’t valid for another week.
Even if you make everything else accept the time, timestamps will be screwed. Like in a fake gmail service that accepts mocked times, what do you use for timestamps on emails?
jmtulloss · · focus · HN ↗
the_mitsuhiko · · focus · HN ↗
That greatly depends on your agent design. If you give a user a full sandbox then you're going to be in a position where you probably need to snapshot it. But there are plenty of agent designs that are not using full VMs and for those the state story is way easier.
deng_nan · · focus · HN ↗
[dead]