What building Tengri taught me about VM sandboxing

I want a coding agent to have a real Linux environment. It needs compilers, package managers, a terminal, an editor, and permission to change the machine when the task requires it. A repository can bring arbitrary install scripts and dependencies with it. The environment has to contain what those programs do.
That is the problem behind Tengri, the agent desktop I have been building in
proompteng/lab. The difficult parts have
been deciding what each guest owns, keeping authority outside that guest, and
making sleep preserve useful work without leaving a VM consuming memory.
The VM provides a kernel boundary. The surrounding system still has to decide who can reach it, which disks it can write, and what happens after a failed restore.
These notes describe the committed Kata implementation at
808bfa8460and the prepared Firecracker work in PR #14813, researched on October 7, 2026. The prepared runtime and its isolated test results are distinct from a production cutover. The authenticated one-second latency target remains unproven.
A worktree answers a different isolation question
One earlier investigation started with a practical question. When several agents use the same shell service, does each one get its own sandbox?
In the setup we inspected, repository sessions allocated separate Git worktrees
and branches inside one shared container. That prevented accidental edits to
the same checkout. The agents still shared a Linux user, home directory,
credentials, network, and compute. An agentId labeled the execution. It did
not create another security boundary.
The current repository session implementation creates those worktrees, and the command runner requires a session for execution. A session makes ownership and working directory explicit. It does not give the process a private kernel.
Different mechanisms answer different questions:
| Mechanism | What it separates | What still needs attention |
|---|---|---|
| Git worktree | Checkout and branch | Processes, credentials, home, and network |
| Separate ordinary container | Process and filesystem views, configured resources | Shared host kernel, mounts, and injected authority |
| Separate VM | Guest kernel and guest memory | VMM confinement, devices, storage, network, and external authority |
Kata's architecture adds hardware virtualization around container workloads. The container's namespaces and cgroups live inside a guest kernel. That extra boundary matters when I want an agent to administer its own environment.
Guest root is an intentional product choice
Tengri's committed runtime creates a MicroVM resource and projects it into a
kata-fc Pod. The fixed guest profile is four vCPUs, 8 GiB of memory, and a
16 GiB persistent home volume. The Go process inside the guest, Nanoagent,
owns files, terminals, editor startup, and the Codex app server. A Rust control
plane owns Kubernetes lifecycle operations.
Tengri allocates one VM per GitHub owner. Terminals and agent processes inside that VM share its filesystem and guest identity. The VM boundary follows the owner, so a new terminal does not create another sandbox.
The request path is browser, authenticated Next.js backend, Tengri, then Nanoagent. The browser does not get a Kubernetes client or a direct guest control credential. The control-plane documentation describes the ownership and transport checks at those boundaries.
Inside the VM, the guest user can become root with passwordless sudo. The
guest has a writable root filesystem and the capabilities needed for Linux
administration. That freedom is useful for a development environment.
The retained home survives a Kata Pod recreation. Running processes, terminal sessions, and changes to the system root do not. Preserving files and preserving a running machine are different lifecycle promises.
The Pod builder
keeps privileged: false, selects kata-fc, disables ordinary service-account
token mounting, and excludes host namespaces and host-path mounts. Those
choices depend on the VM runtime actually providing the intended boundary.
Copying the permissive guest security context into an ordinary container would
change the security model.
Guest root can inspect and change its own guest. That includes credentials available inside it. I cannot promise a secret is hidden from the agent while also giving the agent root on the machine that holds it.
Identity needs a scope smaller than the cluster
The committed control path uses SPIRE mutual TLS and pins the expected peer identity. Signed request metadata carries the authenticated user's identity with replay protection. Nanoagent's workload identity includes the current Pod UID, so a replacement Pod is a different incarnation.
There is an awkward but necessary detail here. A host SPIRE agent cannot attest Unix processes through the guest kernel. The Kata implementation therefore runs a SPIRE agent inside each guest. A guest administrator can access that guest's workload identity. Exact peer checks and Pod-bound attestation limit what that identity can claim. Its private key is not a secret from guest root. The guest identity contract makes that limit explicit.
Authentication also leaves an ownership question. The committed SpiceDB integration checks workspace access against the permission service. Open streams recheck access and close on revocation or authority failure. This is source behavior, not a claim that every deployed guest has been tested against it.
The same separation applies to external accounts. Permission to compile a project inside a VM does not establish permission to use a connector or a cloud account. I want those credentials held by a service outside the guest, with authorization for the requested action. The sandbox should not inherit every credential the platform operator possesses.
Network access needs its own policy too. The committed guest network policy permits public egress while excluding private, link-local, and other reserved ranges, with narrow DNS and identity-service exceptions. That lets development tools download packages. It also means data inside the guest can leave through permitted internet connections. VM isolation does not stop that by itself.
Readiness has to include the first useful operation
The lifecycle problem became visible when creation and wake took too long. Starting a Firecracker process is only part of starting an agent environment. Scheduling, storage attachment, tool installation, editor setup, identity, and app-server initialization can all sit between a request and useful work.
Tengri's repeat-start work validates installation receipts and reuses tools on the persistent home. A cache hit avoids starting Homebrew or Neovim just to prove they were installed. That removes repeated setup. Recreating a Kata Pod still leaves infrastructure and guest startup on the request path.
The next design uses six prepared Firecracker slots. Each slot boots and initializes its own guest before becoming available, commits a private snapshot, and stops the VMM. Creation claims a prepared slot. Resume restores the same owner's last committed state. An exhausted pool returns a capacity error instead of declaring an environment ready while it installs tools.
In the prepared runtime design, readiness means a file read, a terminal round trip, and an already initialized Codex app server work. The target measures the authenticated backend request through those operations at p95. A listening socket is too early to stop the clock.
A snapshot keeps ownership and failure state too
Prepared capacity creates a temptation to clone one initialized guest for everyone. Memory can contain credentials, random state, tokens, and open connections. Firecracker's snapshot documentation puts snapshot management and uniqueness responsibilities on the integrator.
Tengri's proposed runtime gives each slot its own prepared snapshot and disks. It never clones a user's memory into another owner's VM. A lease epoch and durable journal bind operations to the current owner and incarnation. Deletion retires that owner's disks, snapshot, and identity before the slot can serve another owner.
Sleep also needs a precise meaning. Pausing vCPUs retains memory. To release resident guest memory, the runtime must quiesce writes, save and sync a new snapshot generation, commit its journal, terminate the VMM, and evict clean snapshot pages after the mapping closes. Only then can it report sleep complete.
Resident memory and Kubernetes reservations are separate measurements. Stopping the VMM can release guest RAM while a stable slot Pod still reserves resources for its next wake. Capacity accounting has to include both.
The memory snapshot and writable disks must agree. Once the guest writes to its disks, an older memory file is not a safe recovery point for those newer disks. A failed save needs an explicit failure path, such as thawing and resuming the still-live guest.
Loss of contact is another distinct state. A timeout does not prove the old VM has stopped writing. Starting a replacement against the same home requires evidence that the old writer cannot execute or write again. The design retains the claim when that evidence is missing.
Moving SPIRE credentials to a supervisor outside guest memory also avoids restoring host credentials that expired while the guest slept. Guest connections reopen after restore, and readiness checks them again.
The VMM needs containment of its own
The guest boundary depends on a host process. Firecracker's production guidance calls for its jailer, alongside seccomp, namespaces, cgroups, and dropped privileges. Running a microVM does not remove the need to constrain the VMM.
The prepared Tengri implementation explores using the OCI container's isolation around a non-root Firecracker child, with no capabilities and the default seccomp filter. It separates the identity-holding supervisor from the VMM runner and initializes TAP in the Pod's private network namespace. That is a specific experimental launch contract. It needs its own validation and review; I am not treating it as automatically equivalent to the upstream jailer recommendation.
Continuity passed while the latency target failed
The October 7 isolated KVM test report
is the most useful evidence so far. The fixture had one CPU and a 9 GiB limit,
separate network and PID namespaces, and no host data mounts. The 50-cycle run
used baseline revision f7144f0e7e.
| Measurement | Observed result |
|---|---|
| Prepared creation, one sample | 1.919 seconds |
| Cold resume p50, 50 samples | 1.886 seconds |
| Cold resume p95, 50 samples | 3.224 seconds |
| Worst cold resume | 3.829 seconds |
| Sleep and durable snapshot | 47.194 to 177.289 seconds |
All 50 cycles preserved files, the same running shell process, and initialized Codex. Completed sleeps passed the checks for no running VMM and zero resident snapshot pages. Those are useful continuity and memory-release results. The subsecond target failed.
A later single diagnostic on revision 6224276e1 measured creation at
1.751 seconds, cold resume at 1.601 seconds, and sleep at 68.908 seconds. The
runner loaded in about 86 milliseconds, then spent about 692 milliseconds in
the guest resume and readiness hook. A separate Codex request took about
511 milliseconds in the proxy. Those measurements identify stages to
investigate. They do not establish the cause or a new p95 distribution.
The fixture excludes backend authentication, Kubernetes latency, raw PVC allocation, and six concurrent guests. It cannot establish the full product's latency. Passing continuity checks also does not establish complete security acceptance of the runtime.
I like this result because it gives me something specific to work on. Useful state survived, guest memory was released, and the user still waits too long. The next work is to explain those slow stages and measure the authenticated path on the intended guest profile. I want to keep the isolation boundary while making the first file read, terminal command, and agent request fast enough to feel immediate.