Est.

Sandbox Cold Start Latency and Its Impact on Agent Throughput

Invisible sandbox delays compound into hard throughput ceilings as agent concurrency scales.

Contributing Editor · · 11 min read
Cover illustration for “Sandbox Cold Start Latency and Its Impact on Agent Throughput”
Sandbox Infrastructure · September 30, 2026 · 11 min read · 2,380 words

Sandbox cold start latency is the delay between an agent requesting a compute environment and that environment being ready to run code, and it is quietly capping how much work agent infrastructure can push through, even at a few hundred milliseconds a pop. As agent concurrency and task volume climb, that overhead turns into a hard ceiling on throughput because the system can no longer absorb it as a rounding error. Understanding where the time actually goes, and how isolation choices set the floor for it, is the starting point for building agent pipelines that hold up under real production load.

Agent throughput problems that look like something other than infrastructure problems at first

Most teams still think about agent throughput as a simple multiplication: LLM speed times the number of agents running in parallel. Under that model, if tasks are completing slowly, the fix is a faster model or a bigger rate limit.

What teams actually see tells a different story. Queue depths grow, tasks finish slower than the math predicts at scale, and CI pipelines back up in ways that look like rate limiting or model latency but aren't. As of January 2026, 57% of companies run agents in production, putting most organizations past the prototype stage, yet few have instrumented their pipelines closely enough to separate sandbox overhead from everything else eating into wall time Mike Mason. The overhead is there. It's just invisible until someone goes looking for it.

Stripe offers a useful sense of scale here: its agents generate roughly 1,300 pull requests a week Nuvox AI. At that volume, even a small per-task delay compounds into something that shapes how the whole pipeline performs, not a line item anyone can shrug off Nuvox AI. Dropbox made a related observation at DX Annual 2026: once code generation sped up, the bottleneck didn't disappear, it moved downstream, piling pressure onto review queues, CI systems, and validation workflows. The same shift happens when sandbox startup is the slow link in the chain. Speeding up one stage doesn't remove the constraint, it just relocates it. The rest of this piece traces exactly where that overhead originates, how it stacks up as concurrency rises, and which architectural choices actually bring it down.

Sandbox cold start latency and where the time goes

Cold start, defined precisely, is the elapsed time from an agent triggering a sandbox to that sandbox being ready to accept a tool call or execute code. It is not one event, but a sequence: image pull and layer decompression, kernel or hypervisor initialization (which varies heavily by isolation type), process and runtime startup, dependency import, and, if the task requires it, repo cloning or environment hydration.

The AgentCgroup study found that OS-level work, meaning tool calls and container initialization, already accounts for 56 to 74% of total end-to-end latency. Agent logic and model inference, the parts everyone assumes are slow, are not the bottleneck. The dependency import phase alone deserves attention here. Libraries like Pandas, NumPy, PyArrow, and Parquet routinely add hundreds of milliseconds on their own, and stacking agent framework imports (LangGraph, CrewAI) and model client setup on top stretches that window further, all before a single line of task logic runs.

A distinction matters for everything that follows: cold start, booting from zero, is a fundamentally different event from warm resume, returning to a paused or snapshotted sandbox. The two have different latency profiles and call for different fixes, a point the rest of this piece returns to repeatedly. It also matters which kind of coding agent is asking. A PR review agent, a CI triage agent, and a feature-building agent carry different session lengths and hit cold start costs at different rates, so a fix tuned for one won't necessarily suit another.

How isolation technology determines the cold start baseline

Security strength and startup speed move in opposite directions, and whichever isolation mechanism a team picks sets the floor for how fast its sandboxes can start. There's no way around that tradeoff, only ways to manage it.

Three isolation families dominate the landscape.

GPU cold starts fall into a much worse category. Infrastructure provisioning alone can stretch to roughly 19 seconds end to end when scaling from zero on Cloud Run, a real concern for any agent running inference or fine-tuning inside its sandbox Blaxel.

None of this is optional overhead to be engineered away without consequence. Skipping proper sandboxing lets a compromised or malicious tool call inherit the agent process's full permissions: broad filesystem access, environment variables holding API keys, network reach into internal services. The cold start cost buys something real. Teams cannot just chase the fastest isolation number without first asking what their threat model actually requires, because the latency figure only means something next to what that isolation level is protecting. Standard containers (Docker) represent the weakest isolation, sharing a kernel, and their cold starts of 1–5 seconds make them unsuitable for real-time agent interactions. gVisor, which uses user-space kernel interception, offers a middle ground of broad compatibility and reasonable cold starts, and is the isolation technology used by Modal Cosmonic. Firecracker and Kata microVMs provide hardware-enforced kernel-per-workload isolation, the strongest security boundary, but add 150ms–2s to cold starts, and are used by E2B, Blaxel, Northflank, and Vercel Sandbox Nuvox AI Fast.io Cosmonic Blog — Northflank. V8 isolates achieve sub-50ms starts but offer limited language support, as used by Cloudflare Dynamic Workers Nuvox AI Fast.io.

Cold start latency compounding with agent concurrency and task volume at scale

Agents rarely make a single tool call per task. They make many, in sequence, and if each one requires a cold sandbox, startup latency doesn't just add up, it multiplies across every step of the task.

Concurrency makes this worse. Cold start overhead chips away at both of those constraints at once.

Billing compounds the problem further. Traditional cloud billing assumes workloads stay active, but agent workloads burst briefly and then go dormant. Teams end up choosing between keeping sandboxes running and eating the idle cost, or accepting a cold start penalty every single time, and neither option is efficient once volume climbs.

This tends to appear late in metrics, which is part of what makes it dangerous. Productivity metrics like PRs merged or tasks completed look fine early on, and the throughput ceiling only becomes visible once concurrency climbs and queue depths start piling up, by which point the architectural choices behind them are much harder to unwind. At Stripe's scale, with dozens of agents running in parallel across codebases, even a 150ms cold start per tool call aggregates into seconds of wasted wall time per agent session, per hour Fast.io Cosmonic Blog — Northflank. A well-configured agent can process 50–200 tasks per day depending on task complexity and LLM latency, with the practical ceiling being API rate limits and compute budget rather than agent capability, and cold start overhead erodes both. Stated precisely, a sandbox that boots cold rather than resuming in under 25ms carries that delta across every tool call, every concurrent user, and every billing cycle, making the sandbox layer the dominant variable in production performance arXiv Blaxel.

The mitigation toolkit: warm pools, snapshots, and perpetual standby

Four patterns account for most of how teams claw this latency back.

Pre-warmed sandbox pools keep a set of sandboxes already started, repo pulled, dependencies installed, server running, before any user is waiting on one. For predictable load, this drops perceived cold start close to zero. Pre-warming scales with peak capacity projections, so idle compute cost follows right along with it, and under spiky or unpredictable traffic, warm pools either under-serve or overspend. Modal documents this as an explicit pattern built on its SDK primitives.

Filesystem and memory snapshots take a different approach: capture a sandbox's initialized state, dependencies loaded, server running, and restore from that snapshot instead of re-initializing from scratch. Modal offers filesystem snapshots as generally available, directory snapshots in beta, and memory snapshots in alpha, with directory snapshots attachable to pre-warmed sandboxes mid-run arXiv. Memory snapshotting in particular targets initialization-heavy workloads, capturing CPU or GPU memory state to skip past the most expensive part of startup. Research system DeltaBox achieves checkpoint in 14ms and rollback in 5ms by tracking only changed state rather than duplicating the full sandbox, letting agents doing tree search or reinforcement learning explore substantially more nodes inside a fixed time budget arXiv.

Perpetual standby is a third route: rather than cold-booting or keeping a sandbox fully active the whole time, some platforms hold it in a minimal-resource standby state built for fast resume. This is the direct answer to the billing mismatch described earlier: workloads that burst and go dormant stop paying full active rates without eating the full cold start penalty either.

The fourth pattern is architectural rather than a feature: decoupling model inference from tool execution. Extracting inference into its own GPU-backed service, while retrieval and tool execution run on separate CPU-backed services that scale independently, is what makes targeted optimization possible in the first place. Bundle inference and tool execution on the same compute, and none of the warm-standby or pre-warming tricks above can be applied to tool execution specifically. Blaxel achieves a 25ms resume from standby, designed for intermittent, burst-pattern workloads arXiv. Together Code Sandbox offers a 500ms snapshot resume Blog — Northflank.

Platform cold start profiles compared: what the benchmarks show

Comparing platforms on cold start alone misses the point. Isolation technology, cold start baseline, warm resume capability, session duration limits, concurrency ceiling, and GPU support all matter together, because they're the variables that actually determine throughput at scale.

Daytona defaults to Docker, with Kata or Sysbox available for stronger isolation, and posts a headline sub-90 millisecond cold start with unlimited session duration, though pricing runs enterprise-only and the Docker default means isolation is weaker unless Kata is explicitly turned on Northflank. Vercel Sandbox uses Firecracker microVMs with sub-second cold starts, session limits tiered by plan (45 minutes on Hobby, up to 24 hours on Pro and Enterprise), and native Python and Node.js support but no GPU. Cloudflare Sandboxes run Linux containers with sub-50 millisecond cold starts and configurable idle sleep, defaulting to 10 minutes but extendable indefinitely with keepAlive, each sandbox getting its own filesystem, network, and process space across Cloudflare's network, at the cost of narrower language support than the VM-based options.

Depot runs Docker isolation with roughly a 5-second cold start, an 8-hour maximum session, and tiered monthly pricing of $20 (Developer) or $200 (Startup) billed per second above plan allowances, with no GPU. Beam Cloud, also on Docker, posts roughly a 3-second cold start but unlimited session duration and GPU availability, billed by usage.

Nothing here is free. And at the scale end, research on running agents across 10,000-plus concurrent sandboxes on Kubernetes and hybrid cloud makes clear that startup latency has to stay low enough that provisioning itself doesn't become the rate-limiter, which is exactly the regime platforms advertising 50,000-plus concurrent sessions are engineering toward arXiv. Northflank supports Kata Containers (with Cloud Hypervisor as primary VMM), Firecracker, and gVisor, allowing a choice of isolation per workload, with a 97ms median TTI and 167ms burst per the July 2026 ComputeSDK benchmark, unlimited session duration, BYOC to AWS, GCP, or Azure, processing of over 2 million isolated workloads monthly, GPU options including L4, A100 40GB/80GB, H100, and H200, and usage-based per-second billing, as stated in Northflank's blog. E2B uses Firecracker microVMs with a 150ms cold start, a 24-hour maximum session limit requiring checkpointing for longer tasks, no GPU support, no BYOC, a stated usage by 88% of Fortune 100 companies for frontier agentic workflows, and a free hobby tier plus paid plans Modal. Blaxel uses Firecracker microVMs with a 25ms resume from standby under a perpetual standby model, unlimited session duration, no GPU per the sourced comparison table, contact sales pricing, and is designed specifically for intermittent, burst-pattern agent workloads. Together Code Sandbox uses microVM isolation with a 500ms snapshot resume, VM-style pricing, and is positioned for Together AI users. The key pattern across the table is that the fastest cold starts (Cloudflare sub-50ms, Blaxel's 25ms resume, and Daytona's sub-90ms) each involve a tradeoff of limited language support, weaker default isolation, or burst-resume rather than true cold boot Northflank. The research-grade ceiling is set by the DeltaBox system (ArXiv), which achieves a 14ms checkpoint and 5ms rollback, not yet a commercial product but a signal of where OS-level optimization is heading for stateful agent workloads.

Governance and Observability Requirements in the Cold Start Equation

None of the numbers above happen in isolation from governance work, and that's where a lot of published benchmarks quietly understate the real cost. Session-scoped credentials, minted fresh per session and revoked on completion, audit logging of every tool call and diff, and hard budget caps all add initialization work, and most of that work happens during or just after the sandbox boots.

A sandbox that starts in 90ms with no credential injection is not the same as one that starts in 90ms and also mints scoped credentials, attaches the right repository permissions, and opens an audit log stream, so teams should benchmark their actual initialization path, not vendor baselines. Compliance posture is part of the same calculation. Modal carries SOC 2 Type II certification with HIPAA support on its Enterprise plan, and Northflank and E2B carry their own compliance postures, and for regulated industries that overhead isn't optional, it has to be weighed against raw latency rather than traded off against it.

Observability carries its own tax too. Full logging of tool calls, diffs, and thinking tokens adds I/O overhead per session, and at tens of thousands of concurrent sessions that overhead becomes a real infrastructure line item, not an afterthought Modal. Defining agents as version-controlled configuration, YAML checked straight into the repo, means a sandbox's initialization becomes deterministic and auditable, since the same config governing what an agent does also governs how its sandbox gets provisioned. And enforcing hard per-session budget caps requires the sandbox layer to emit spend data in real time, something platforms without per-second billing struggle to support cleanly. The billing model is central to the cold start conversation. It's part of the same architecture decision, made at the same time, by the same team.

Sources

  1. Best Code Execution Sandboxes for AI Agents in 2026 | Modal Blog
  2. 10 Best Code Execution Sandboxes for AI Agents (2026) | Fastio
  3. What’s the best code execution sandbox for AI agents in 2026? | Blog — Northflank
  4. Top AI sandbox platforms in 2026, ranked | Blog — Northflank
  5. AI Sandbox: The Complete Guide to Sandboxing AI Agents in 2026 | Cosmonic
  6. Best Code Execution Sandboxes for AI Agents in 2026 | Blaxel
  7. AI Coding Agents in 2026: Coherence Through Orchestration, Not Autonomy | Mike Mason
  8. DeltaBox: Scaling Stateful AI Agents with Millisecond-Level ...

More in Sandbox Infrastructure