Sandboxing an Agent That Executes Code
The threat model is unusual and that is what makes it easy to get wrong. The code is not written by an attacker, and it is not written by a trusted developer either. It is written by a model that has been reading attacker-controlled text all afternoon. What you are actually defending against Three sources, ordered by how often they bite: Accident. An rm -rf with a variable that was empty, a…
We haven't written up this one. Dev.to has the full story — the link below goes straight to it.



