Network Egress Controls for AI Coding Agent Sandboxes
Stopping compromised agents from reaching the network requires more than sandboxes.

Network Egress Controls for AI Coding Agent Sandboxes.
Network egress controls, not just sandboxes, for AI coding agents
Network egress controls are not a bonus feature bolted onto an AI coding agent's sandbox. They decide whether a compromised agent stays contained or gets to talk to the outside world. The distinction matters more now than it did even a year ago, because coding agents no longer just autocomplete a function or suggest a diff. They read an entire repository, plan changes across dozens of files, write the code, run the test suite, open the pull request, and in some setups merge it without a human ever looking closely. That's a fundamentally different risk profile than a completion tool sitting in an editor.
The gap between piloting these agents and actually running them in production is wide. Something like 88% of agent pilots never make it to production, and the research points at infrastructure, not capability, as the reason: isolation, governance, and compliance controls that security teams demand before an agent touches anything that matters augmentcode.com. Ask security leaders what's holding back agentic AI at scale and security concerns top the list, cited by nearly two-thirds of respondents in one 2026 trust survey, with 72% naming cybersecurity specifically as a serious risk augmentcode.com.
Part of the reason is that the threat model here doesn't resemble the one security teams built their playbooks around. A traditional application follows a fixed set of code paths that a static policy can reason about in advance. An agent generates its actions at runtime, often from inputs an attacker can shape, so the API calls it decides to make and the resources it decides to touch aren't knowable ahead of time. OWASP AIVSS assigns a CVSS v4.0 Base Score of 9.4 to the scenario where an agent is manipulated into executing attacker-provided arbitrary code.
The clearest illustration of what this looks like in the wild came from the so-called ZombAIs incident. During what looked like a routine browsing task, Claude landed on a malicious page, picked up an injected prompt, downloaded a binary, ran it, and connected out to a command-and-control server. The whole chain worked on the first attempt, and there was no sandbox in place to stop any of it.
That's the case for treating egress as its own discipline rather than an afterthought to filesystem isolation. A sandbox that locks down disk access but leaves the network wide open blocks some categories of damage and does almost nothing against this one. Egress control would have broken the ZombAIs chain at the download step or the callback step, and it produces the requirements the rest of this piece works through: default-deny policy, domain allowlisting, credential isolation, static IPs, and the isolation substrate that supports them all. None of these are optional add-ons. Each closes a different gap, and skipping one leaves the others doing less than they're supposed to.
Default-deny egress in a sandbox as the necessary starting point
Network egress locked to default-deny, filesystem boundaries, process isolation, secrets that stay scoped, and a lifecycle that doesn't persist state across sessions unless someone explicitly asks it to are the five things a sandbox worth the name enforces. Egress comes first on that list for a reason. Get it wrong and the blast radius of a compromised agent stops being local, it goes global.
Default-deny means every outbound connection gets blocked unless it's on an explicit allowlist. The alternative, allowing everything by default and trying to block known-bad destinations, is backwards for agent workloads, because the bad destination in question hasn't been enumerated by anyone yet. It might be a domain an attacker registered yesterday.
The instance metadata service address makes the argument for default-deny about as concretely as it can be made. Agents that can reach 169.254.169.254 (AWS Instance Metadata Service) can acquire the host instance's IAM credentials Blaxel. This is a well-documented lateral movement technique that any competent attacker already knows. It's a well-worn lateral movement technique, and default-deny egress kills it outright, because there's simply no allowlist entry for that address. It doesn't matter what the agent decides to try. The door isn't there.
Platforms building for this already treat default-deny as a starting assumption rather than a configuration option. Buildkite's Cleanroom runs coding workloads in a microVM sandbox with deny-by-default egress paired with a credential proxy, baked into the platform rather than left to whoever configures the deployment. On the container side, the rough equivalent is launching Docker with the --network none flag, though containers share a kernel with the host, which makes that isolation boundary weaker than a microVM's from the outset. A different approach, the ceLLMate system, enforces default-deny at the HTTP layer through a Chrome extension policy engine, and in testing it predicted the correct policy over 94% of the time and stopped all 12 simulated prompt injection attacks thrown at it, at a cost of 7 to 15% added latency Daytona.
None of that means default-deny solves the whole problem by itself. An agent that needs to pip install a package, hit a third-party API, or clone a private repo has to be able to do those things, so default-deny is the floor a system builds on, not the ceiling. The next layer, the allowlist, is what makes the floor livable. One design principle carries through everything that follows: the agent itself should never be able to touch its own egress policy. If it can edit the rules that govern what it's allowed to reach, the rules aren't really rules. The IMDS attack vector stands as the clearest argument for default-deny.
Domain allowlists and their limits under real agent workloads
An allowlist enumerates the specific hostnames or CIDR ranges the agent legitimately needs; everything else is blocked at the egress layer. For a typical coding agent that list tends to include any third-party APIs the agent is explicitly authorized to use and dependency CDNs used by the build toolchain.
Daytona's implementation shows how this gets built in practice: domain allowlists with limits on entry count, wildcard support, and an option to block the network. Credentials get handled through placeholders that only resolve to real values on outbound requests matching an allowlisted host, so a request to anywhere else keeps the placeholder and never sees the actual secret.
Where this starts to strain is under the kind of workloads agents actually run. Wildcard entries like *.amazonaws.com are often necessary but dramatically expand the allowed surface, since S3, STS, and the IMDS all resolve under that wildcard. Allowlisting a package registry doesn't vet what gets pulled from it either. pypi.org being on the list says nothing about whether the specific package the agent decides to install that day is safe.
Agents doing research or browsing tasks present a harder case still, because their entire job is to reach destinations nobody enumerated in advance. An allowlist model built around known hostnames doesn't have much to say about a workload whose purpose is discovering unknown ones. And even a well-built allowlist can be undermined by DNS rebinding: a domain that's legitimately on the list can resolve to an internal IP address if DNS resolution itself isn't locked down alongside the domain policy.
Where the allowlist gets enforced changes how much any of this matters. Enforcement at the kernel level, through eBPF or at a TAP device (the pattern some platforms call CubeVS uses per-sandbox TAP devices with eBPF policy checks), is much harder for a compromised agent to route around than enforcement living inside the same container process the agent runs in. And the list itself needs to come from a human and live in version control, not get written or edited by the agent at runtime. That's the same immutable-config principle from the default-deny layer, applied one level up. The LLM API endpoint being called.
Credential isolation at the network layer: why secrets must never enter the sandbox environment
The most common mistake in agent sandbox design is injecting API keys and cloud credentials into the sandbox as plain environment variables. Any process running inside that sandbox, including a hijacked agent or a package quietly poisoned somewhere upstream in its supply chain, can read those variables and ship them out. There's no exploit required for this to go wrong. The agent reads something it shouldn't, it already holds a valid credential, and it makes an outbound call that's technically permitted. Nothing breaks the rules; the rules just weren't strict enough.
The fix is architectural: hold secrets outside the sandbox entirely and inject them at the network layer, on the outbound request itself, so a plaintext key never sits inside the environment the agent's code runs in. Cloudflare's Sandboxes product documents exactly this pattern, injecting secrets through an egress proxy so they never appear in the agent's code or environment Cloudflare. Blaxel does something similar: its proxy intercepts outbound HTTPS traffic and injects headers, body fields, and secrets server-side, meaning plaintext secrets never make it into the sandbox at all. Daytona's version uses the placeholder mechanism described above, substituting the real value only when the destination matches an allowlisted host, and leaving the placeholder untouched for anything else.
Run the ZombAIs scenario again with this control in place. An agent hijacked through prompt injection can still try to exfiltrate data to a command-and-control server, but it has nothing to hand over, because the API keys never existed inside its environment in the first place. The attacker gets whatever the agent can see and touch. It doesn't get the keys to anything else.
Credential scoping adds another layer on top of this. Minting a fresh, narrowly scoped credential per session and revoking it the moment the session ends means a credential stolen from a finished session is already dead by the time anyone could misuse it. This runs alongside proxy-based injection rather than replacing it. And for regulated environments, this isn't just good practice, it's often the requirement itself: frameworks like SR 11-7 and HIPAA call for credentials and sensitive data to never transit an untrusted execution environment, and proxy-based injection satisfies that condition in a way environment-variable injection simply cannot.
Static egress IPs for enterprise firewall allowlisting
If a sandbox vendor is asked which static IPs a security team should put on its firewall allowlist, on most platforms today the honest answer is that there isn't one, not without the customer standing up and operating their own proxy VM. That's a procurement problem as much as a technical one, and it appears the moment an enterprise security review gets past the first round of questions.
Static egress IPs matter for reasons that go beyond firewall configuration. A lot of third-party APIs, payment processors, internal enterprise services, certain SaaS platforms, restrict access by source IP on top of whatever credential gets presented, and a dynamic IP pool means an agent's requests get blocked regardless of whether its credentials check out. Incident response depends on knowing which IP an outbound connection actually came from, and that gets murky fast behind a shared dynamic NAT pool. Rate limiting on external APIs frequently keys off source IP too, which means one tenant's misbehaving agent can burn through a shared pool's rate limit and take other tenants' traffic down with it.
Coverage across the sandbox market is uneven. Blaxel has dedicated egress gateways in private preview, giving sandboxes allowlistable static outbound IPs, listed as a platform primitive. Fly.io offers inexpensive static egress IPs alongside network policies, a fit for teams already building their own orchestration layer on top. Modal gates static IPs behind its Team plan, priced at $250 a month Blaxel.
That same roundup shows that static egress IPs, custom domains, proxy-based secrets injection, and egress filtering aren't edge-case asks from unusually paranoid customers, they're routine procurement requirements. That reframes the whole conversation. This isn't a security nice-to-have anymore, it's table stakes for anyone trying to sell into an enterprise. And a static IP alone doesn't finish the job either. An IP shared across every customer on a platform still can't give one specific enterprise customer something it can put on its own firewall allowlist without also allowlisting every other tenant on that platform. Dedicated egress gateways, not shared NAT, are what actually deliver on the promise. Static egress IPs matter for reasons that extend beyond firewall allowlisting. The Blaxel roundup assesses platform coverage on static egress IPs. Daytona has no confirmed static egress IP support as of the roundup's publication. E2B requires self-hosted IP tunneling through a gateway VM for static IPs.
The isolation substrate egress controls run on top of
None of the controls described so far mean much if the isolation layer underneath them can be broken. gVisor's own documentation states that with a standard container, the workload sits one system call away from compromising the host. And because a container that escapes shares the host's network namespace, it also escapes whatever egress policy was enforced at the container level. The policy doesn't survive the breakout.
This isn't theoretical. CVE-2024-21626, known as Leaky Vessels, showed that a crafted Dockerfile WORKDIR directive in versions of runc up to 1.1.11 could resolve to the host filesystem, a container escape that took the egress policy down with it once achieved. CVE-2025-59528 in Flowise scored a full 10.0 on CVSS, with roughly 15,000 instances exposed, demonstrating that shared-kernel isolation is insufficient for production AI agents handling untrusted code gVisor Blaxel.
Three isolation tiers exist, and they carry different implications for egress. MicroVMs, the Firecracker and Kata Containers family, give each workload its own dedicated kernel with hardware-level isolation, and the TAP network interface is the natural chokepoint for egress policy: policy enforced there survives a container-level escape, because there's no shared kernel to escape into.
gVisor takes a different route, running a user-space kernel that intercepts syscalls through a component called Sentry before they ever reach the real kernel Modal. Escaping it requires a bug in Sentry and a bug in the host kernel, both, which makes it stronger than a bare container but weaker than a microVM gVisor Modal. It underlies Cloud Run, App Engine, and Cloud Functions, with some I/O overhead of 10–30%, and Modal uses it, with VM Sandboxes still in alpha gVisor. Hardened containers reduce capabilities and add seccomp profiles or AppArmor and SELinux policies, but remain the weakest isolation boundary, appropriate for trusted code in development but not for untrusted AI-generated code in production.
The eBPF and TAP enforcement pattern mentioned earlier, per-sandbox TAP devices paired with eBPF policy checks, pushes egress enforcement down to the kernel level, below both the container and the VM guest, which makes it materially harder for a compromised agent to route around. It's described in the research as becoming the standard shape for agent network isolation generally.
Isolation tier isn't a platform-wide decision so much as a per-workload one. Northflank, for instance, handles more than 2 million isolated workloads a month and supports four separate isolation technologies, Firecracker, Kata Containers, gVisor, and Cloud Hypervisor, letting a team pick the tier that fits a given workload rather than being locked into one for everything. GPU workloads complicate that choice further: gVisor's user-space interception blocks direct PCIe passthrough for GPU calls, while Firecracker's hardware virtualization path supports VFIO passthrough, which forces a real tradeoff between isolation strength and GPU access for any coding agent running GPU-accelerated builds or tests.
A study covering five AI sandbox products found that the different engine classes, microVM, userspace kernel, and OCI container, separate cleanly from each other on nearly every architectural measure, but products sitting within the same class don't separate cleanly from one another. The paper from arXiv:2606.08433 found that downstream pin policy is the dominant operator-facing variable, with patch lag ranging from 0 days to 471+ days to "opaque" to effectively infinite for some products. That has a direct bearing on egress security. Choosing a microVM substrate and then running it on stale runc or Firecracker binaries quietly gives back a good chunk of the isolation advantage that substrate was supposed to provide. Patch cadence is part of the egress security posture itself, and it deserves the same scrutiny as the allowlist and the credential proxy sitting on top of it. It's part of the egress security posture itself, and it deserves the same scrutiny as the allowlist and the credential proxy sitting on top of it. Firecracker powers Lambda, Fargate, Fly.io, Vercel Sandbox, and E2B, with cold starts at 150ms for E2B and Daytona. Blaxel runs microVMs on a custom Firecracker fork, with sandboxes resuming from standby in under 25ms. SOURCE PAGES (what the pages behind the outline's links say).
Sources
- Best Code Execution Sandboxes for AI Agents in 2026 | Blaxel
- AI Code Sandboxes: A Comparative Security Study. Part 1 of 2 – Engine-Level Properties (Attack Surface, Leakage, Stackability, CVE History, Patch Cadence, Fuzzing)
- How to sandbox AI agents in 2026: MicroVMs, gVisor & isolation strategies | Blog — Northflank
- blaxel.ai
- fly.io


