This is an excellent explainer. I noticed one important flaw in the setup, though: Toad assigns multiple puzzles on page 5, but by page 10 all the machines seem to be working on the same puzzle.
I'm not sure how this maps onto reality. If each agent had a different problem to solve, how did they cooperate?
It's complicated, but one of the ways they first cooperated was to break the flag generation algorithm.
I'm sort of winging this, but the ExploitGym challenges involve exploiting a program then getting an HMAC generated "flag" which the scorer can verify. The agents collaborated on a way to produce a legitimate flag for all of their tasks (funny enough this was actually what caused them to escape sandbox, the agents were incorrectly convinced that simply providing a correct flag without accurate steps to exploit the program in their transcripts would cause them to still fail the task)
munchler · · focus · HN ↗
I'm not sure how this maps onto reality. If each agent had a different problem to solve, how did they cooperate?
jabedude · · focus · HN ↗
I'm sort of winging this, but the ExploitGym challenges involve exploiting a program then getting an HMAC generated "flag" which the scorer can verify. The agents collaborated on a way to produce a legitimate flag for all of their tasks (funny enough this was actually what caused them to escape sandbox, the agents were incorrectly convinced that simply providing a correct flag without accurate steps to exploit the program in their transcripts would cause them to still fail the task)