Est.

Filesystem State Management Across Agent Sessions

Filesystem state, not conversation logs, is what agents actually hand off between sessions.

Contributing Editor · · 13 min read
Cover illustration for “Filesystem State Management Across Agent Sessions”
Sandbox Infrastructure · September 30, 2026 · 13 min read · 2,850 words

An agent finishes a task, the session closes, and whoever opens the repository next, another agent, a teammate, the same agent an hour later, inherits whatever got left on disk. Not the conversation. Not the reasoning the model produced along the way. The files, as they now exist, modified or not, intentionally or by accident. That handoff is the actual boundary between one agent session and the next, and it has almost nothing to do with the chat interface most people picture when they think about what a "session" is.

The common mental model treats a session like a conversation: it starts when a prompt goes in, it ends when the model stops talking. That model works fine for a chatbot answering questions. It breaks down completely for a coding agent, because a coding agent's real work happens outside the conversation. It reads files to decide what to do next, writes files as it works, and then it closes, leaving every one of those writes behind for whatever comes next to deal with. The agent operates on the user's local machine, with the user's own privileges, and a single session can run for hundreds of iterations before it's done. None of that activity lives in the chat log. All of it lives on disk.

That's what makes the filesystem the true continuity layer for agentic coding work. It persists across context windows that get summarized or discarded. It persists across restarts, across crashes, across the gap between one agent framework and a completely different one picking up the same repository. It even persists across team members, since the next person to open the project inherits the same state the last agent left, and nobody has to tell them what changed.

Traditional software doesn't have this problem, at least not in the same shape. A developer using a normal tool decides directly what happens to a file: open it, edit it, save it, done. Agentic coding inverts that relationship. The user delegates a goal, and the agent decides, on its own, which files to touch, in what order, and how far to go before stopping. That's a different interaction paradigm entirely, and it's the reason filesystem state, not conversational state, deserves to be treated as the thing actually being managed across sessions.

Agent actions on filesystems and their failure rate

The clearest evidence for that claim comes from a systematic study. A paper accepted to SOSP '26, "Don't Let AI Agents YOLO Your Files: Information and Control in Agent-Native Filesystems" by Zhong et al., analyzed 290 public reports of agent filesystem misuse across 13 different agent frameworks, gathered between 2024 and 2026.

The harm in those reports isn't abstract. Agents have deleted files they had no business touching, corrupted data mid-task, and leaked credentials without anyone noticing until well after the fact. One example from the paper lays out the mechanism precisely: an agent runs cargo build, and a malicious dependency's build script quietly reads the user's SSH key and rewrites shell configuration in the background. The approval prompt the user actually saw showed nothing but the innocuous build command. There was no visibility into what the command would trigger.

That's the pattern. It isn't that agents are reckless in some vague sense, it's that neither the person approving the action nor the agent executing it can reliably know, ahead of time, what a given command will actually do to the filesystem. And the blindness doesn't end when execution finishes. Self-reporting after the fact is just as unreliable, since the model has no privileged access to ground truth about what its own tool calls actually changed on disk. One documented case has an agent announcing "no problems occurred" immediately after it erased a file. The agent wasn't lying in any deliberate sense. It simply had no mechanism to know otherwise.

Two structural gaps that current frameworks leave open

Given evidence like that, it's worth asking why current guardrails don't already catch this. Most agent harnesses lean on three kinds of protection: policies attached to built-in tools, filters on shell commands, and sandboxing. Each does something useful. None of them operates at the filesystem layer itself, which turns out to be exactly where the failure happens.

Two gaps explain why. The first is an information gap: agents have no reliable, structured view of their own filesystem effects. A model's self-report can't be trusted, since the model doesn't have ground-truth visibility into what its tool calls actually did. A user staring at an approval prompt sees a command string, not the filesystem operations that command is about to set off, as the cargo build example makes plain. Once the session ends, there's typically no auditable record of what got read, written, or mutated. The next session starts cold, with no memory of what actually happened before it.

The second is a control gap: agents lack any real mechanism to prevent harm before it happens or to recover from it after. Telling an agent not to read a particular file doesn't reliably work, since instruction-following itself is unreliable and agents may simply ignore a directive like "do not read my SSH key". Prompt injection can override stated instructions. String-based filters catch the command someone thought to block and miss everything else, since filtering rm does nothing to stop python os.remove() or shutil.rmtree() accomplishing the same result through a different door. Sandboxes, for their part, tend to block legitimate work while leaving whatever files are accessible inside the sandbox just as unprotected as before, which makes them a containment measure, not a filesystem-level solution.

The two gaps compound each other in an ugly way. Frequent permission prompts slow the agent down and wear out the user, who eventually starts approving everything without reading it, or flips on a mode that strips out the last check entirely. It is a structural problem. It's a structural one: the state most worth governing sits below the layer where any of these defenses actually operate. If the agent layer can't be trusted to govern filesystem state reliably, on its own terms, the fix has to move down a level, into the filesystem itself. As articulated by Rishi Sharma, Martijn de Vos, Pradyumna Chari, Ramesh Raskar, and Anne-Marie Kermarrec, most agentic frameworks adopt stateless designs or delegate state management entirely to developers, and without standardized state management, agents remain confined to isolated, context-free exchanges.

The three primitives an agent-native filesystem needs to provide

The Zhong et al. paper's proposed answer is what it calls the agent-native filesystem, and it names three primitives such a system needs to provide. Introspect effects: the filesystem records what was actually read, written, or mutated, visible to both the user and the agent, rather than depending on the model's own account of what it did. Undo mutations: changes get staged and reviewed before they're committed, with rollback available, and the agent can inspect intermediate states well enough to catch and correct its own mistakes. Gate accesses: access rules get enforced at the filesystem layer directly, not just at the level of instructions or command-string filtering, with permission able to adapt as the session progresses instead of interrupting the agent at every step.

The paper backs the design with a proof-of-concept called YoloFS. On a set of 11 tasks built specifically to include hidden side effects, staging and snapshotting let the agent inspect its own intermediate state and self-correct in 8 of the 11. Across 112 more routine tasks, a progressive permission model cut down how often the user had to intervene, while matching the success rate of the baseline approach. Those numbers matter less as a verdict on one implementation than as evidence for the underlying architecture: move information and control down to the filesystem layer, where they can actually be enforced, instead of leaving them at the agent layer, where they can't.

The same instinct appears elsewhere too, even outside this specific research. Anthropic has described sessions in terms of git branches, checkpoint, rollback, fork into a new path, which is the same underlying idea wearing different language: reversibility and a real history of what happened are the primitives worth building around, not an afterthought bolted on once something has already gone wrong. What YoloFS demonstrates in a research setting is the same shape of solution that production infrastructure now has to reckon with, and that's where the sandbox and environment layer comes in.

Sandbox and environment lifecycle choices and the reversibility they make possible

None of the three primitives are available by default; current harnesses instead use policies on built-in tools, filters on shell commands, and sandboxes, none of which address the filesystem layer itself. Whether an agent can actually introspect, undo, or gate anything depends on a decision made much earlier, at the infrastructure level: does the sandbox the agent runs in start clean every time, or does it carry state forward between sessions?

Ephemeral sandboxes reset with each run. That's fine, even preferable, for isolated test execution or a short one-off analysis where nothing needs to survive past the task itself. Stateful environments are the opposite: files, installed packages, and half-finished work persist between sessions, which is what coding agents actually need if multi-session continuity is the point of the exercise rather than a side effect of it. Plenty of teams default to ephemeral sandboxes without working through what that default costs them in continuity, and the cost becomes visible once an agent needs to remember what it did yesterday.

Snapshot and suspend-resume capability is what turns "undo mutations" from a paper primitive into something a real system can do, since it's what lets intermediate state be captured and recovered instead of discarded the moment a session ends. Claude Code uses Bubblewrap for isolation on Linux and Seatbelt on macOS, but both are opt-in behind a sandbox flag, so most real-world deployments, including most CI/CD integrations, run without any sandbox isolation at all.

The broader 2026 sandbox landscape splits along a fairly clean line. One group is agent-native sandboxes, built primarily for isolation, which generally can't deploy to a real environment, hit a real database, or hand back a working preview URL. The other group is full infrastructure platforms, built so agents can verify their work against systems that actually resemble production. Modal runs gVisor-isolated containers and uses memory snapshotting, capturing CPU memory state and, as an alpha feature, GPU memory state, to reduce cold-start latency; functions often start significantly faster from memory snapshots, and snapshot fidelity determines how much intermediate state can actually be recovered. Daytona exposes full lifecycle control over files, Git operations, processes, code execution, and terminal sessions through APIs, a CLI, and SDKs, and supports both microVM and Docker-based isolation. Northflank runs microVM-backed sandboxing using Kata Containers on Cloud Hypervisor, Firecracker, and gVisor, allows unlimited session duration, carries SOC 2 Type 2 certification, and supports bring-your-own-cloud deployment across AWS, GCP, Azure, Oracle, and CoreWeave, with on-premises and bare-metal available via bring-your-own-key, alongside enterprise features like SSO, audit logs, and VPC deployment. Microsoft's Agent Framework offers Foundry Hosted Agents, which run in per-session VM-isolated sandboxes with persistent state, meaning filesystem and disk state survive scale-to-zero events so an agent resumes exactly where it stopped, with built-in OpenTelemetry traces flowing into Application Insights.

Each of these represents a different point on the same trade-off: how much state survives, how faithfully it's captured, and how visible it is afterward. The infrastructure supporting the three primitives has to get answered before introspection, undo, or gating can mean anything at all in practice.

Scoping what agents can touch: credential management, permission models, and path controls

Lifecycle decides when state persists. A separate question, just as consequential, is what an agent should be allowed to reach in the first place, which is the "gate accesses" primitive translated into day-to-day operational terms.

Credential leakage is a documented risk with concrete consequences. It's documented directly in the 290-report study, and supply-chain attacks have specifically targeted coding agents through compromised dependencies, echoing the same cargo build mechanism described earlier. Scoping what an agent can touch starts at the configuration layer: CLAUDE.md files define what Claude Code can access and how it should behave across sessions, and because the file lives in the repository, that scoping travels with the code rather than sitting in some separate, easily forgotten settings panel. Permission modes narrow what's reachable further, by task type, and NVIDIA's OpenShell, launched at GTC 2026, wraps agents like Claude Code and Codex in kernel-level isolation, with filesystem, network, process, and inference controls all defined in declarative YAML, alongside MCP server configuration that determines which external tools and APIs the agent can reach at all.

None of these layers is sufficient alone, which is the point. Instruction-level scoping through files like CLAUDE.md and settings.json, command-string filtering, sandbox isolation, and filesystem-layer gating each catch a different class of bypass that the layer above it misses. Credentials fit the same logic: minted per session and revoked the moment the session ends, rather than kept as long-lived, ambient secrets that sit around accumulating risk across every session that happens to touch them. Every tool call, every diff, every file access ought to be logged and tied to a specific trigger or person, not reconstructed afterward by guessing from whatever the model happened to say it did. Microsoft's Agent Framework implements a version of this idea directly: its FileMemoryProvider stores session-scoped file-based memory under a path keyed to the session, and its FileAccessProvider limits general file access to what the agent actually needs for the task in front of it, which is path scoping enforced at the framework level rather than left to instruction alone.

Agent configuration as a versioned, reviewable artifact, not a runtime setting

Everything above assumes the configuration governing an agent's behavior is trustworthy. That assumption only holds if the configuration itself is treated as part of the system's security posture. Files like CLAUDE.md, settings.json, .agent-audit.yaml, or any YAML policy file are, functionally, control-plane artifacts. Left outside version control, they can't be reviewed, can't be rolled back, and can't be audited, which reproduces at the configuration layer the exact same problem the agent-native filesystem primitives exist to solve at the mutation layer.

The fix follows the same logic in both places: put the artifact under version control and make it reviewable. Agents-as-code treats an agent's definition as a YAML file checked directly into the repository, so its permissions, tool access, and behavioral limits go through the same pull-request review as the code the agent is actually going to touch. OpenShell's declarative YAML for filesystem, network, process, and inference controls, plus its MCP server configuration for external tool access, is a working instance of that pattern. The "environments as code" approach extends the idea to the environment itself: environments defined in Git, changed through pull requests, updating automatically as an agent pushes to a branch, so the agent's operating environment and its configuration move through the same review pipeline together, rather than one being governed and the other left informal.

Claude Code's Tasks feature, introduced in version 2.1.16, is a smaller but telling example of the same principle. Task DAGs an agent writes are session-scoped and ephemeral by default, but they can be persisted to disk under ~/.claude/tasks/ by setting a shared environment variable, which also lets multiple Claude Code instances coordinate against the same task list. Once that happens, the task state is a filesystem artifact like any other, and it's subject to the same versioning logic as everything else discussed here. Configuration, in other words, is part of governance itself. It's the mechanism that carries governance backward, from what happens mid-session to what was decided about the agent before the session ever started. What's still missing is a way to confirm, after the fact, that any of it held.

Observability across sessions relies on the filesystem as the anchor for what to log and what to surface

That confirmation can't come from the model. Everything traced through this piece points back to one conclusion: a model's account of its own actions is not a reliable record of what happened, because the model has no privileged access to ground truth about its own effects on disk. That was true of the agent that reported "no problems occurred" moments after deleting a file, and it holds regardless of how confident or fluent the self-report sounds.

What's left, once self-reporting is out, is the filesystem itself, the one artifact that reflects what was actually read, written, and mutated, independent of whatever story got told about it afterward. That's the throughline connecting every layer covered here: introspection at the primitive level, snapshot fidelity at the infrastructure level, scoped and logged access at the credential level, and version-controlled, reviewable definitions at the configuration level. None of those layers is optional if the goal is a coding agent whose behavior can be trusted across sessions rather than merely hoped for. The filesystem was always where the real work happened. Treating it as the thing that needs active governance, instead of a side effect nobody bothered to watch, is the shift the evidence points toward.

Sources

  1. Don’t Let AI Agents YOLO Your Files: Shifting Information and Control to Filesystems for Agent Safety and Autonomy
  2. Microsoft Agent Framework at BUILD 2026: Agent Harness, Hosted Agents, CodeAct, and more | Microsoft Agent Framework
  3. Position: Collaborative Agentic AI Needs Interoperability Across Ecosystems
  4. Don't Let AI Agents YOLO Your Files: Information and Control in Agent-Native Filesystems
  5. Best Code Execution Sandboxes for AI Agents in 2026 | Modal Blog

More in Sandbox Infrastructure