kodebeat / themes / ai-infrastructure

What did your coding agent just try to run?

Most teams adopted AI coding tools faster than they could govern them.

Containment, then visibility

The boundary comes first, then the record: put a limit around the agent, then get a record of what happened inside it. Anything not publicly released is marked, because "I built it" and "you can go and read it" are different claims.

Containment

An agent should reach one project directory and the network you allowed, and nothing else. Each of these takes one axis away rather than trying to be a policy engine.

jails a coding agent to one project azkaban severs a process's network offline scores a command's blast radius before it runs scoville

The gateway

Every model call through one endpoint, so quotas, redaction and an audit log are a configuration change rather than a rewrite.

OpenAI- and Anthropic-compatible proxy with rotation and failover chicco which models this machine can actually run llm-fit

Verifying the output

Code an agent wrote still has to be owned by a human. These check it mechanically instead of hoping the review catches it.

secret scanning across tree and history gandalf grades a repo against a policy rubric gradebook multi-pass LLM audit that tracks what was never read readthrough cross-validated PRDs and quality gates ai-harness in development

The papers

Every paper carries the methodology behind any number it emits, and what the tool deliberately does not do. Pick them up from the catalogue.

  • Containing what a coding agent can reachazkaban and offline, two Linux sandboxes with one axis each
  • Pooling several model providers behind one local endpointa local OpenAI- and Anthropic-compatible rotation proxy
  • Deciding which models a machine can run before downloading onea hardware-aware model, quantization and runtime advisor
  • Auditing code with a model, and knowing what was never reada multi-pass LLM code audit with coverage accounting
  • Surviving a session limit without losing the sessiona PTY wrapper that waits out the session limit and resumes

Agent containment is the part teams are hitting right now with almost nothing available to buy. If that is where you are, the papers describe the approach in enough detail to build it yourself.

Working on this?

Tell me what you are looking at and I will tell you honestly whether any of this helps. No pitch attached, and the papers are free either way.

Send a request

Send a request

One message, no follow-up sequence.

Your address is used for this request and nothing else.