What did your coding agent just try to run?
Most teams adopted AI coding tools faster than they could govern them.
Containment, then visibility
The boundary comes first, then the record: put a limit around the agent, then get a record of what happened inside it. Anything not publicly released is marked, because "I built it" and "you can go and read it" are different claims.
Containment
An agent should reach one project directory and the network you allowed, and nothing else. Each of these takes one axis away rather than trying to be a policy engine.
The gateway
Every model call through one endpoint, so quotas, redaction and an audit log are a configuration change rather than a rewrite.
Verifying the output
Code an agent wrote still has to be owned by a human. These check it mechanically instead of hoping the review catches it.
The papers
Every paper carries the methodology behind any number it emits, and what the tool deliberately does not do. Pick them up from the catalogue.
- Containing what a coding agent can reachazkaban and offline, two Linux sandboxes with one axis each
- Pooling several model providers behind one local endpointa local OpenAI- and Anthropic-compatible rotation proxy
- Deciding which models a machine can run before downloading onea hardware-aware model, quantization and runtime advisor
- Auditing code with a model, and knowing what was never reada multi-pass LLM code audit with coverage accounting
- Surviving a session limit without losing the sessiona PTY wrapper that waits out the session limit and resumes
Agent containment is the part teams are hitting right now with almost nothing available to buy. If that is where you are, the papers describe the approach in enough detail to build it yourself.
Working on this?
Tell me what you are looking at and I will tell you honestly whether any of this helps. No pitch attached, and the papers are free either way.