Building Anvil: code-as-action with a capability sandbox that explains its refusals.
The agent wrote import socket . Now what? Every agent framework got very good at making models write code. Almost none got good at the question that follows immediately afterwards: what is that code allowed to do to my machine? I built Anvil to answer that in the smallest honest shape I could: an agent writes a Python tool for a task over some real CSVs, the tool executes in a locked-down…
Anvil is a code-as-action system that utilizes a capability sandbox to explain its refusals. The system is designed to ensure that the code generated by an agent is only allowed to perform specific actions on a machine. The core functionality of Anvil consists of three main components: the generated code, the sandbox trace, and the artifact. These components are presented in a three-pane interface, with an additional panel for testing break attempts.
The system consists of four layers: the AST pre-check, the sandboxed subprocess, the artifact collection, and the trace of every refusal. The first layer, the AST pre-check, involves parsing the candidate source code without executing it. This layer checks for banned imports, process primitives, dynamic eval, and any literal paths that resolve outside the allowed roots.
Each rule generates a reason string that is displayed verbatim in the user interface and asserted by the test suite. For example, if the import of the socket module is blocked due to policy rules, the reason "network denied by policy" will be displayed. Similarly, if an attempt is made to write outside of the allowed paths, the reason "outside ALLOWED_WRITE_PATHS" will be shown.
The second layer involves the sandboxed subprocess. The tool executes in a locked-down subprocess, where every attempt to escape is blocked and reported with the exact reason string, at the layer that detected the attempt. Layer 3 deals with the kernel, where the system proves that the sandbox is effectively enforcing its rules, rather than simply assuming that it is.
The system uses the Seatbelt kernel sandbox, which provides network denial and write confinement enforced below Python. This ensures that even if a C extension attempts to read sensitive files like /etc/passwd, the system will block the read operation.
Throughout the system's development, several lessons were learned. For instance, over-eager static rules can lead to false positives, and banning introspection modules can unintentionally block legitimate tools. The system also discovered that some audit events lack a caller frame, requiring special handling. Additionally, the system had to remove certain legitimate reasons for denial due to technical limitations.
Overall, Anvil provides a robust and transparent framework for executing code generated by agents, while clearly explaining any refusals that occur.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.