The methodology of the whole experiment feels unfair. It's almost as if they tweaked the experimental setup to give them the results they wanted.
Choosing and older model and a permissive prompt, and explicitly no HITL for confirmation, is what got it to “12% of runs leaked data”.
Lots of comments here talk about using a bad model, but beyond that, why did a shopping agent have an unrestricted browser and a memory containing card details, date of birth, and SSN information? A review promising a discount is exactly the kind of task relevant bait an agent will encounter. Better models may follow it less often, but I wouldn’t want the payment and privacy boundary to depend entirely on the model recognizing it.
Disclosure: I’m building Sangria, which lets agents discover products and buy with prepaid credits and spending limits. If this is something that sounds interesting, would love your feedback on the product and the direction we're taking (getsangria.com)
jaredee · · focus · HN ↗
Choosing and older model and a permissive prompt, and explicitly no HITL for confirmation, is what got it to “12% of runs leaked data”.
Lots of comments here talk about using a bad model, but beyond that, why did a shopping agent have an unrestricted browser and a memory containing card details, date of birth, and SSN information? A review promising a discount is exactly the kind of task relevant bait an agent will encounter. Better models may follow it less often, but I wouldn’t want the payment and privacy boundary to depend entirely on the model recognizing it.
Disclosure: I’m building Sangria, which lets agents discover products and buy with prepaid credits and spending limits. If this is something that sounds interesting, would love your feedback on the product and the direction we're taking (getsangria.com)