Sandbox Infra · Three-Way Teardown

How DSec, WeEnv and AgentENV Actually Engineer Sandboxes

Three institutions, three released systems, one shared diagnosis: the environment, not the agent, is where the time goes. A side-by-side teardown of DeepSeek's DSec, Tencent WeChat's WeEnv and Moonshot's AgentENV.

DSec · arXiv:2609.22978• WeEnv · arXiv:2609.30766• AgentENV · open source• EN + 中文•
English 简体中文
Part 3 of 3 The three-way teardown — three released systems, read against the same measurements
Contents
  1. III. DeepSeek / DSec: Turning One Platform into Four Isolation Backends, Then Colluding with the Trainer
  2. IV. Tencent WeChat / WeEnv: Managing the Entire Chain from Packaging → Initialization → Provisioning
  3. V. Moonshot / AgentENV: Turning "an Existing Image" into "a Snapshotable, Forkable, Pausable Environment"
  4. VI. Horizontal Comparison: Three Institutions' Three Answer Sheets
  5. VII. The Only Cross-Institution Measurement: WeEnv's Bench, and Four Qualifications That Must Be Added
  6. VIII. What Each of the Three Institutions Left That Can Be Directly Taken Away
  7. IX. Closing
ProjectInstitutionMaterial FormTime
DSec (DeepSeek Elastic Compute)DeepSeek-AI + Tsinghua UniversityTechnical report, arXiv:2609.22978v1 [cs.DC], 31 pages2026-09-19
WeEnvTencent WeChat AI (WeChat AI, Tencent)Paper, arXiv:2609.30766v1 [cs.DC], 16 pages2026-09-25
AgentENV (AENV)Moonshot AIOpen-source codebase, GitHub org kvcache-aiFirst open-source commit 2026-07-25

The problem these three projects aim to solve has its best name in WeEnv: environment tax.

2.1 WeEnv Quantified the "Tax"

The WeChat team used their production-grade framework Slime to break down iteration time (WeEnv §1, §3.1, Figure 2):

BackendTrainEnv initTask execOthers
E2B (microVM)24.0%53.4%20.2%2.5%
Docker (container)15.4%39.5%44.0%1.1%

Original wording:

"initialization alone accounts for 53.4% of the iteration time with virtual machines (E2B) and 39.5% with containers (Docker), 2.6× the task execution itself with E2B and roughly on par with it under Docker, while training takes merely 15.4%–24.0%."

On E2B, environment initialization eats 53.4% of iteration time—2.6× task execution and more than 2× training itself. Training accounts for only 24.0%.

Another perspective: rollout (= initialization + task execution) accounts for 73.6% on E2B and 83.5% on Docker. In other words, the vast majority of this pipeline's time is not spent learning, nor solving problems, but preparing the place to solve problems.

2.2 DSec Measured the Same Disease from the Production Side, but with a Different Set of Metrics

DSec did not do an "iteration time share" breakdown. It measured the shape of the workload (§4, data sampled from one week in early 2026):

Bursty creation. Number of sandboxes created per task (Fig. 2):

Sandboxes created per taskp50p90p99
Containers2,5287,96916,388
microVMs3521,8354,044

"The largest production jobs can request up to 32K sandboxes. These requests arrive within a short window because the training or evaluation batch cannot use an instance until its environment is ready."

Extremely sparse access. Actual runtime access ratio of container images by language (Tab. 3):

Image typeC++GoJavaJavaScriptPython
Fraction of data accessed8.7%13.3%9.2%4.2%6.0%
Image size4.9 GB4.1 GB12.1 GB9.6 GB6.0 GB

The vast majority of image data is never touched from beginning to end.

Low fanout. Image reuse within a single task (Fig. 8, over 1.5 million containers / 390,000 microVMs): container images p50 = 3, p90 = 28; microVM images p50 = 1, p90 = 3. DSec's conclusion is blunt: the working set is too diverse for single-node caching to absorb, so every burst necessarily triggers image pulls.

Sparse CPU + long-lived memory. Fig. 5: about 90% of container and microVM sandboxes use on average only within 5% of their requested CPU—a natural justification for oversubscription. Fig. 7: container lifetime p50 = 17.4 min, p99 = 231.5 min; microVM p50 = 15.5 min, p99 = 213.9 min. CPU is idle, but memory and state remain pinned.

2.3 Two Independent Instruments Measured the Same Set of Conclusions

Putting the two sets of measurements side by side reveals something very convincing—two companies, without knowing about each other, observed the same set of structural facts:

Structural factDeepSeek (DSec §4)WeChat (WeEnv §3)
Sparse accessOnly 4.2%–13.3% of image data touched at runtimeMedian reads only 0.81% of artifact, p99 only 6.02%, no task exceeds 7%
Bursty creationUp to 32K sandboxes per task; container p99 = 16,388A batch is admitted in two waves; initialization is on the critical path
Long-tail CPU demand~90% of sandboxes average ≤5% of requested CPUPeak CPU p50 = 0.99 cores → p99 = 68.3 cores, nearly 70× span
Long-tail memory demandMemory reclamation is the density bottleneck92% of tasks peak < 0.2 GiB, but p99 jumps to 7.0 GiB
Dramatic demand fluctuationsCPU intermittent during tool calls, memory persistently residentSingle-task CPU demand surges 47× within one second

This table is the most important argumentative foundation of the entire article. It shows that the "environment tax" is not a defect of any particular implementation, but a structural property of the agentic RL workload. Therefore, the disagreement among the three is not about "whether to solve it," but about "which cut to make first."

2.4 And the Three Cut at Completely Different Places

This is the skeleton of the entire article. Here is the conclusion first:

                     Common disease: environment tax
              (bursty creation · sparse access · long-tail demand · state residency)
                            │
        ┌───────────────────┼───────────────────┐
        ▼                   ▼                   ▼
   DeepSeek/DSec       WeChat/WeEnv      Moonshot/AgentENV
   "Platform + density" "Full lifecycle"   "Startup + state"
        │                   │                   │
   Bet: one platform     Bet: manage the      Bet: turn "an existing
   manages four isolation entire chain from   image" into "a running
   backends, coordinates  packaging to        environment," and make
   with RL trainer,       provisioning        it snapshotable,
   handles cluster                          forkable, pausable
   elasticity
        │                   │                   │
   Scarce resource:      Scarce resource:      Scarce resource:
   node density &        iteration time        startup latency &
   cluster elasticity    (the tax itself)      state portability
        │                   │                   │
   Evidence: 160 nodes/  Evidence: init share  Evidence: README claims
   3M sandboxes/day      53.4% → 9.1%          <50ms / <100ms / 9.6×

The following three sections dissect each of the three narrow gates one by one.


III. DeepSeek / DSec: Turning One Platform into Four Isolation Backends, Then Colluding with the Trainer

3.1 Positioning

DSec is a technical report (not a paper submission), authored by DeepSeek-AI and Tsinghua University. First author Jialiang Huang is a Tsinghua PhD student who completed the work during an internship at DeepSeek-AI; corresponding author is Liyue Zhang. It serves the sandbox workload for all RL training and evaluation from DeepSeek V3.2 to V4.1 (§6 opening).

Its self-positioning is very clear, and it deliberately avoids a single abstraction:

"This interface is intentionally not a full semantic abstraction over all backends. Function calls, containers, microVMs, and full VMs have different startup costs, isolation boundaries, filesystem semantics, and operating-system capabilities. libdsec provides a unified access path and a similar operational model, but the caller remains responsible for selecting a backend that matches the workload." (§2.1)

This is the only design among the three that "does not bet on a single isolation mechanism."

3.2 Production Scale: These Numbers Define All Subsequent Design

DSec §2.4 gives actual numbers for a scale unit:

MetricValue
Nodes~160 CPU nodes
Cores30K cores
DRAM~250 TB
Images/layersSeveral PB
Daily sandbox instances~3 million
Peak concurrency~380,000
Creation rate>5,000 instances/sec

And the density operating point (§4.3): production stably runs 3,200 containers or 800 microVMs per node (one-day sampled peak was 1,048 containers / 524 microVMs).

This magnitude explains the aggressiveness of all DSec design. When the creation rate is 5,000/s, "pulling the full image" is not slow—it is a structural error.

3.3 Architecture: Control Plane, Runtime, Storage—Three Layers

DSec §3's architecture is divided into three parts, each deliberately made stateless and replaceable:

Cluster-level services (control plane)

Node runtime

An easily overlooked but informative design choice: FnCall and containers both run inside QEMU/libvirt VMs, rather than directly on the host (§3.3).

"FnCall and containers run inside QEMU/libvirt VMs rather than directly on the host. The VM provides an isolated kernel and network stack and serves as an additional security boundary between untrusted containers and the bare metal."

That is, a VM wrapped around the container—sacrificing density for "untrusted containers do not directly touch bare metal." This "rather have one more layer" orientation reappears in §6.4's security story.

3.4 Mechanism One: Composable Environment Layers—Turning O(m·N) into O(m)

Problem (§4.2): A sandbox's content can be decomposed into three parts—base image (OS-level dependencies), workspace (task code repository + task-specific dependencies), toolkit (frequently updated tools, e.g., DeepSeek Harness).

Single-week production data (Tab. 2):

Backendbase imagesworkspacessnapshotsAggregate size
Containers11,266102,171–82.8 TB
microVMs253,5904,88950.9 TB

There are also 103 toolkits, and 67.8% of sandboxes need a workspace or toolkit beyond the base image.

If they were welded into a single OCI image: maintaining M bases, N workspaces, K toolkits, upgrading m bases requires rebuilding O(m·N), upgrading k toolkits requires O(k·N). The original example is intuitive: upgrading Toolkit T1, every monolithic image containing it must be rebuilt, even if not a single byte of base or workspace has changed.

DSec's answer (§5.1): Treat the three as logical layers with independent lifecycles, dynamically stacking them with overlayfs at creation time. Base at the bottom, workspace inserted above as a read-only layer, toolkit on top.

The "dirtiest" and most interesting implementation detail: modifying dockerd.

"We modify the open-source Docker daemon (based on the Moby project) to dynamically insert EROFS-backed lower layers at container creation time. Specifically, we pass the path of a pre-mounted EROFS layer and insert it into the overlayfs stack before mounting, placing it as the topmost lower layer so it can override files in layers below. This change is minimal, requiring only 30 lines of Go code." (§7)

30 lines of Go exchange O(m·N)/O(k·N) → O(m)/O(k). This is the highest cost-performance engineering change in the entire article.

Why EROFS instead of ext4/XFS: EROFS is designed for read-only data, eliminating write-related accounting, with a more compact layout, supporting compression while preserving random access. Key comparison:

"Unlike a tar.gz archive, EROFS can read and decompress only the compressed blocks covering the requested data, so the complete image does not have to be transferred and unpacked before use."

Measured gains (§8.3): Same workspace + toolkit, tar.gz per-sandbox extraction vs EROFS direct mount:

SchemeEnd-to-end completion timeDisk write trafficPeak write throughput
tar.gz extraction79 min~5.5×~3.4×
EROFS mount45 min (1.76× speedup)1×1×

Note the original's explanation of "EROFS peak CPU is higher"—not because overhead is greater, but because more sandboxes enter the tool-calling phase concurrently earlier. This self-clarification is very honest.

The microVM side gets the same composition model (§5.1 end): base and toolkit are packaged as independent versioned EROFS images, exposed to the guest as read-only block devices; inside the guest, rootfs uses overlayfs, with the EROFS mount point as lower and a directory on the ext4 writable disk as upper.

3.5 Mechanism Two: On-Demand Loading—Writes Stay Local, Reads Go Through 3FS in Batches

DSec's key judgment is "do not build a separate image distribution layer" (§5.3):

"Existing on-demand image distribution systems often combine a container registry with peer-to-peer delivery to prevent the registry from becoming a bottleneck. We instead host images on 3FS, which already supports our production training workloads at scale. This choice reuses the existing storage infrastructure and avoids deploying a separate image-distribution layer."

But 3FS has a characteristic that must be accommodated: very high large sequential read/write throughput, very poor small random I/O. This asymmetry directly determines three design principles:

  1. Writes stay local—sandbox writes are uncontrollable (logs and other small frequent writes), so the writable layer is placed on node-local disks, completely bypassing 3FS's small-write penalty.
  2. Reads are on-demand and batched—read-only data is fetched from 3FS only when accessed, and fetched in batches to saturate 3FS's large I/O throughput.
  3. Metadata as local as possible—metadata is often accessed by small reads; prefetch to local when separable.

Container path uses three EROFS features:

It also does something to "reduce mount count": continuously collapsing layers within a size threshold (~3 GB) offline into a pair of metadata/data EROFS images, while preserving overlayfs whiteout semantics to correctly express file deletion.

The microVM path is different, and the reason is hard (§5.3):

"Docker's overlay2 driver cannot use an overlayfs-backed data directory. An alternative would be to export the host-mounted filesystem to the guest via virtio-fs, but our Firecracker backend does not support this interface."

So microVMs go block-level: read-only base/toolkit layers still use EROFS, while the writable ext4 disk goes through OverlayBD, exposed via ublk. The cost is stated openly: "Unlike the multi-device EROFS path, ext4 metadata remains embedded in the block image, so metadata reads can trigger remote I/O." Mitigation is 256 KiB chunk fetching + local second-level filesystem cache—after a chunk is evicted from page cache, it remains in the local cache and does not need to go back to 3FS.

Measured gains (§8.2): 10 nodes, 8,192 containers, real evaluation workload:

SchemeCompletion timeCumulative disk writes
Docker cold pull (eager)>60 min>1,600 GB/node
EROFS on-demand~35 min~700 GB (57% less)
Docker full local cache (upper bound)~35 min~600 GB

On-demand loading matches the theoretical upper bound of "full local precaching", while being 1.71× faster than eager and writing 57% less. This result is more meaningful than "how much faster"—it shows the on-demand path introduces no additional penalty.

3.6 Mechanism Three: Memory and CPU QoS Under Oversubscription—Two Truly Hard Engineering Conclusions

This section is the most "only known from production experience" part of DSec.

Memory (§5.2 / §8.4). microVMs have two independent sources of memory waste:

  1. Page cache duplication: image data read through virtual block devices is cached once by the host and once by each guest.
  2. Guest idle pages are not returned: idle pages inside the guest are not proactively reported, so the host cannot reclaim them; and requested capacity often far exceeds actual demand, so the guest has no internal pressure to reclaim inactive pages. Lifetimes are long (p50 15.5 min, p99 > 3 hours), amplifying residency cost.

Two complementary answers:

The cost is also stated clearly: pmem requires the guest to allocate struct page for the entire range; with 4 KiB pages + 64-byte struct page, metadata consumes 1/64 of pmem capacity—"a 128 GB pmem device requires 2 GB of guest RAM for this metadata." And cold access may trigger synchronous page faults.

The two mechanisms are complementary: pmem eliminates duplication, DAMON+FPR reclaims idle pages. After combination, instantaneous peak CPU rose from 26.5% to 41.4%—DSec does not hide this and gives operational advice: "In CPU-constrained deployments, operators may prefer to enable FPR alone and retain virtio-blk."

CPU QoS (§5.2 / §8.5). The problem scenario is specific: some tasks have strict per-step latency budgets (e.g., board-game agents with fixed per-step time limits), and "lowering best-effort priority" is simply not enough—because BE and LS may run on the two SMT sibling threads of the same physical core, still competing for execution resources.

Two-layer strategy:

  1. Put BE into SCHED_IDLE, yielding CPU when LS is runnable;
  2. Enable Linux core scheduling for LS, preventing unrelated BE work from running on sibling threads of the same physical core.

Measured (§8.5, LS per-step latency under 50% BE load):

ConfigurationLatency inflation
No protection baseline45.2%
SCHED_IDLE onlyAt best improves 3.4%
SCHED_IDLE + core scheduling17.3%

"SCHED_IDLE alone improves at best only 3.4%" is a very valuable number—it proves that priority adjustment alone is nearly ineffective against SMT interference, and core scheduling must be touched.

The residual 17.3% is also explained: mainly from turbo frequency reduction under multi-core high load, memory bandwidth and shared LLC contention, which core scheduling cannot address; and it explicitly says memory bandwidth isolation is not done because it is "already tolerable."

3.7 Mechanism Four: Colluding with the RL Trainer—This Is DSec's True Moat

If you remember only one thing about DSec, it should be this: it designs the sandbox platform and the RL trainer as one system. This is something the other two did not do.

(a) Let agents build their own environments (§6.1): The title is "Build environments of Agents, by Agents, for Agents". The core interface is called pack_diff:

"at any point, an agent can checkpoint a sandbox by taking an incremental disk snapshot, which can later be restored as a new sandbox. This checkpoint-and-restore interface turns an interactive session directly into a reusable environment, allowing environments to be built, validated, and consumed on the same infrastructure without a separate image-building pipeline."

Note the positional relationship here: both WeEnv and DSec are solving "environment packaging," but in opposite directions. WeEnv offline decomposes components into layer groups and recomposes them; DSec lets agents interactively freeze a running sandbox into an environment online. The former optimizes "combinatorial explosion of releases," the latter optimizes "there is no build pipeline at all."

Supporting governance measures are equally specific: constraints on agent build rules (limiting performance impact on shared infrastructure), an internal platform for quality inspection and exporting standard formats, builder and runtime using different accounts, and clearing build-time residual data from the writable layer before packaging to avoid bringing reference answers into the image.

(b) Splitting rollout out of the preemptible GPU pool (§6.2): The old version put the agent loop together with model serving and the RL framework in preemptible GPU training pods. When the GPU job is preempted, the agent loop is lost but the sandbox remains, so recovery can only rely on command log replay for reconciliation (reusing recorded results for completed operations, avoiding duplicate side effects of non-idempotent commands).

From DeepSeek-V4.1 onward, rollout execution was moved onto DSec, split into two components:

Both run outside the preemptible GPU pool. The result is that the rollout lifecycle is decoupled from the trainer lifecycle; the worker container + agent sandbox jointly hold the complete rollout state and serve as the single source of truth; preempted GPU jobs do not need command-log replay to rebuild execution after reconnecting.

(c) Pause means reclaim (§6.3): When a GPU job is preempted, the RL framework proactively sends pause to all relevant sandboxes; DSec reclaims memory but retains execution state; any subsequent request to a paused sandbox transparently resumes before execution. The two paths differ:

Backendpauseresume
Containerdocker pause freezes the process tree → enable memory.swap.max → memory.reclaim proactively reclaims anonymous and file pagesMADV_WILLNEED async prefetch on process memory mappings → docker unpause
microVMSnapshot memory and execution state → terminate the Firecracker process to release runtime memoryStart a new process → restore snapshot and continue execution

Note that microVM pause really kills the process, retaining memory via snapshot to disk—this is the source of "pause consumes almost no memory," and is also the contrast point for AgentENV using the same trick later.

3.8 Mechanism Five: Treating Agents as Adversaries—Only DSec Seriously Did This

DSec §6.4 is the most "production-flavored" section of the entire piece, titled "Agent Misbehavior and System Failures". It classifies risks into two categories: the task appears to pass, but the answer comes from an unintended channel (undermining training/evaluation validity), and agent actions damage the environment (endangering this task or other tasks on the same machine).

The failure cases listed in the original are all very specific:

Finding answers inside the sandbox

Finding reference implementations outside the sandbox

Ordinary commands crashing infrastructure

Mitigations (§6.5) only admit to solving part of the problem; the title itself states "These controls address only part of the problem and do not provide a general defense against destructive behavior such as triggering kernel bugs."

Strategic implication of this section: DSec's adversary model is adversarial. WeEnv also uses mechanisms (independent clean evaluation environments against reward hacking), but that is architectural isolation; AgentENV currently has no equivalent security narrative. This difference is expanded in Section V.

3.9 DSec's Costs and Boundaries


IV. Tencent WeChat / WeEnv: Managing the Entire Chain from Packaging → Initialization → Provisioning

4.1 Positioning

WeEnv comes from Tencent WeChat AI (WeChat AI, Tencent), authors Yang Yu, Jing Lei, Shaoxun Zeng (co-first authors, Shaoxun Zeng corresponding), paper arXiv:2609.30766v1, 16 pages. The conclusion is one sentence: "WeEnv is deployed for agentic RL at WeChat."—already online, but no scale numbers disclosed.

Among the three, WeEnv's stance is the most "management school":

"The root cause is the lack of a full-lifecycle solution to environment management." "WeEnv manages environments across packaging, initialization, and provisioning."

It also makes two choices the others did not:

  1. Containers only (WeEnv builds on container tooling: each environment is instantiated as a container, §5). No microVMs, no full VMs.
  2. Introduces a third dimension: provisioning. DSec and AgentENV both treat "resource quota" as a static parameter at creation; only WeEnv treats it as a continuous online scheduling problem.

Plus a methodological uniqueness: it is the only project that actually measured AgentENV as a baseline. This makes it the only source of cross-institution measured data among the three, and also brings fairness issues that must be flagged (see 4.7 and Section VII).

4.2 Three-Part Diagnosis: Packaging Dilemma, Initialization Amplification, Provisioning Rigidity

WeEnv §3's motivation analysis is very clean, with quantification in each of the three parts.

(1) Packaging is a dilemma (§3.2)

SchemeApproachCost
Dynamic assemblyArtifact excludes evolving components, installs at initialization~20 seconds extra per environment to install harness (~13% of initialization)
Static bundlingFull packaging into monolithic artifact, zero install at initializationCombinatorial explosion: T tasks × H harness variants = T×H pre-packaged images

The pain of static bundling is made very clear in the original: "Worse, the harness is not the only fast-evolving component; as described in §2, the evaluator is also updated routinely against newly discovered reward hacking. Each such change touches every affected image, which can take days and lengthens the experimentation cycle of agentic RL."

Note an implicit echo between two institutions here: WeChat says "the evaluator must be frequently updated to combat reward hacking"; DeepSeek uses an entire section (§6.4) to record the reward hacking behaviors they actually caught (forged chronus RPC, log searching, XFS_IOC_SWAPEXT bypassing access control). Two companies independently confirm the same thing: combating reward hacking is a continuous engineering effort requiring frequent environment iteration, not a one-time data cleaning.

(2) Initialization's I/O amplification (§3.3)

WeEnv traced all I/O to the rootfs during each task's execution (on OverlayFS, because partial writes first read the whole file, so measuring only reads covers all required bytes), obtaining a CDF:

"the median task reads merely 0.81% of its artifact, and even the 99th percentile reaches only 6.02%; no task in our trace touches more than 7%." "fetching the full image transfers over 100× the bytes that execution ever accesses."

(3) Provisioning's fixed quota (§3.4)

Profiling peak demand from 8k tasks of SWE-Smith-Python (Figure 4):

Metricp50p90p99Span
Peak CPU0.99 cores17.7 / 18 cores68.3 / 68 cores~70×
Peak memory0.15 GiB0.17 GiB7.0 GiB—

Memory is even more skewed: 92% of tasks peak below 0.2 GiB, barely above the 0.15 GiB median, then p99 jumps directly to 7.0 GiB. The original's arithmetic is striking:

"One sized for the median starves the tail, while one sized for the tail wastes resources on the majority, e.g., a 7 GiB quota would over-provision nine out of ten tasks by more than 35×."

And quota directly translates into latency (Figure 5): for the same resource-intensive task, under the minimum quota (2 cores / 8 GiB) rollout takes 310.5 s, while under the maximum quota (128 cores / 512 GiB) it takes only 201.6 s—tight quotas stretch rollout by 1.5×.

Why prediction is not possible (§4): WeEnv sampled 10 resource-intensive tasks aligned at their start to observe demand evolution (Figure 6). The conclusion is prediction is infeasible—demand has no clear pattern, the same task varies greatly across rollouts, and bursts are sudden: one task's CPU demand surged ~47× within one second. So WeEnv simply gives up prediction and switches to elastic adjustment.

4.3 Mechanism One: Layer Group + Environment Plan—Making Layers First-Class Deployable Units

Problem positioning is very precise (§4):

"The environment tooling of both virtual machines and containers already embraces a similar notion, the layer. ... However, layers serve only as packaging-stage units of caching and deduplication. At initialization, they must have been bound into a single monolithic artifact, one template or one image, before an environment can launch." "The fundamental gap is that layers cannot be located on their own. They exist only as entries in the artifact's manifest... Layers thus cannot be indexed apart from the enclosing artifact."

Key technical detail: Docker's registry protocol only accepts image references, not bare digests. So under the standard protocol, layer-only addressing is unavailable.

WeEnv's answer (§6.1):

  1. Give each layer group its own explicit reference (harness:2.0, task1:1.0, base-env:3.0), independently versioned and published;
  2. Use an environment plan to describe the environment—an ordered list of group references by stacking order;
  3. Each reference still uses the same mechanism to resolve that group's layer digests (query registry for group manifest);
  4. Cache resolution results, so one group reference is resolved only once, not re-resolved at every initialization;
  5. The only new step is concatenation: concatenate the digest lists of each group from bottom to top into a flat layer stack, handed to the mount path;
  6. Order is override semantics: groups higher in the plan override those below;
  7. OverlayFS mounts it as a single read-only rootfs, with an environment-private writable layer on top absorbing all writes.

The original points out the only essential difference from a conventional artifact:

"the artifact freezes its layer stack at packaging time, whereas the plan assembles it at initialization from independently published groups, so the same groups recompose into different environments without repackaging."

And the cost of composition is extremely low: "it manipulates only metadata, resolving cached references, concatenating digest lists, and issuing one mount, while no layer content is copied, unpacked, or executed."

Packaging granularity cut by change rate (§6.2). This is the design guideline itself: "The guideline is to draw group boundaries along change rates." So it cuts into four groups:

GroupContentChange rate
harness groupHarness coordinating model and environment (Claude Code / Codex, etc.)Highest, continuously upgraded
evaluator groupScoring and reward computationHigh, continuously patched against reward hacking
task groupInitial state per task (git checkout, faulty patch)Medium
base groupOS, language runtimes, common toolsLow, shared by all tasks

So a harness upgrade or evaluator patch publishes only a small group, touches none of the T task groups, and the change takes effect in every environment composed afterward.

One more layer: repository-granularity packaging (§6.2 end). The original observes a very strong structure:

"many RL tasks share one code repository and differ only in a small mutation, such as checking out a particular commit or applying a faulty patch. For example, the 39,471 tasks of SWE-Smith-Python span merely 131 repositories."

So task is packaged at repository granularity (all tasks in the same repository share), each task only needs a cheap mutation applied at initialization—measured median 0.43 seconds (1024 initialization measurements), compared to ~20 seconds to install harness.

Results (§9.3, Table 1):

DatasetDynamic assemblyStatic bundlingWeEnv
SWE-Smith-Python39,47139,471·H133 + H
SWE-Smith-C++5,1235,123·H71 + H
SWE-Smith-Go1,6291,629·H21 + H
SWE-Smith-Java6,7046,704·H61 + H
SWE-Smith-JS6,0736,073·H36 + H
SWE-Smith-Rust5,3115,311·H41 + H
SWE-Smith-TS5,0325,032·H32 + H

The general formula is B + 2 + H (B repository-level base images + 1 base tool + 1 evaluator + 1 per harness). From 39,471 (or ×H) down to 133 + H, about two orders of magnitude. More critically, the change cost: under static bundling, changing harness or evaluator requires touching "up to T images of the entire dataset"; under WeEnv, exactly one layer group is republished.

4.4 Mechanism Two: On-Demand Pull + Rebuilding the Layer Format—Because gzip Does Not Allow Seeking

Data flow (§7.1): To intercept file I/O, WeEnv mounts the rootfs into the environment via FUSE. When a task accesses a file, the kernel forwards I/O to envlet's FUSE handler, which uses layer metadata to translate it into a request for "a specified layer + specified byte range." Requests are served in three tiers:

envlet cache (shared by all environments it hosts)
   └─ miss → node-level cache (one per node, shared with peer nodes)
        └─ miss → remote registry

Only reads enter this path—all writes fall to the local writable layer; partial writes to files backed by read-only layers trigger copy-up, which first reads the whole file, so it also appears as a read.

There is also a best-effort background fill: once an environment starts, it traverses the artifact chunk by chunk, skipping chunks already brought in by on-demand fetching, keeping completion off the critical path.

Why the layer format must be rebuilt (§7.2). Two requirements, Docker's conventional layer format satisfies neither:

RequirementConventional format (gzip-compressed tar stream)Consequence
Metadata independently pullableNo central directory; each file header is adjacent to its data, interleaved with contentRecovering metadata requires walking the entire layer
Arbitrary byte ranges independently readablegzip compresses an archive into a single sequential stream; later bytes reference earlier onesReading any byte means decompressing all content before it

WeEnv's layer format:

Measured gains (§9.4):

MetricOn-demand off (full download then start)WeEnv on-demandImprovement
Median initialization32.5 s10.6 s3.1×
P9583.5 s35.1 s2.4×
Worst case208 s (layer must come from remote registry)47 s4.4×
1st RL iteration (all cold)1,845 s763 s2.4×
Sum of all iterations8,092 s5,988 s1.35×

The original's qualitative conclusion is key: "on-demand fetching removes the artifact transfer from the critical path regardless of the cache state"—initialization latency no longer depends on artifact size.

4.5 Mechanism Three: Elastic Provisioning—Three Principles, the Third the Sharpest

Monitoring (§8.1): Every environment starts from the same small quota: 2 cores / 8 GiB. envlet manages resources via cgroup, and monitors through the same interface—reading cpu.stat / memory.stat every 200 ms. It emphasizes that monitoring itself is extremely cheap: nothing runs inside the environment, each environment round takes tens of microseconds, and at 200 ms intervals uses less than 0.1% of a single core.

Principle One: Scale up aggressively, scale down conservatively (§8.2). The reason is that demand can rise 47× in one second, and an incorrect downscale directly kills the environment and loses accumulated multi-turn interactions.

Pressure levelCriterionAction
Normal pressureCPU: quota utilization ≥85%, or 10% of scheduling periods throttled; memory: working set ≥60% of quota, or cgroup reports reclaim events/pressureDouble quota
Severe pressureCPU: throttling exceeds 80%; memory: OOM emergencyQuota immediately ×4
Downscale (CPU)Utilization persistently below 30% for 10 seconds and throttling negligibleHalve quota
Downscale (memory)Working set persistently below 20% of quota for 60 secondsHalve quota

Principle Two: Scale-up must not starve admission (§8.2). This scenario is designed very realistically: because some tasks must be evaluated in independent clean environments, a batch is admitted in two waves—first execution environments, then evaluation environments after execution completes. The two waves overlap on the same node; if scale-up of running environments consumes all admission capacity, evaluation environments cannot start. Two constraints:

Principle Three: Failures must be explicit, not silently degraded (§8.2).

"A trajectory silently corrupted by starvation is worse than a lost one, as it feeds wrong signals into training. When an OOM emergency persists with no tier left to raise, the envlet terminates the task and reports the cause, so that the RL framework re-samples the task instead of learning from a starved execution."

This is the most copy-worthy design judgment in the entire article. It elevates "resource allocation failure" from a performance issue to a training data correctness issue: rather lose a trajectory than let a starved trajectory enter the training set. The same thinking appears again in WeEnv's architecture—evaluation must run in an independent clean environment, "so that rollout-side modifications cannot contaminate the reward computation, guarding against reward hacking."

Measured gains (§9.5):

MetricFixed quotaElastic provisioningImprovement
Single-task rolloutSlowest 2 hit 1,800 s agent timeoutSame task ~40 s45×
Executions saving >100 s—61—
3rd / 5th / 8th iteration1,949 / 1,773 / 1,981 s542 / 806 / 534 sup to 3.7×
Sum of all rollout latencies8,968 s5,988 s1.50×

The original also proactively discloses counterexamples: a few executions on the right end are actually slower under elastic mode, because "their extra time is spent in the solving phase, where the agent takes a different, longer trajectory, a stochastic effect of agent sampling rather than of the resource allocation." This clarification is very well done.

Quota timeline (§9.5, Figure 12) gives a complete example: CPU starts at 2 cores, one burst during solving jumps from 4 directly to 16, then gradually halves back to 2 as the burst subsides (released cores "serve other environments on the same node for most of the solving phase"); during evaluation another burst occurs, quota climbs past 32 up to the single-environment cap 64. Memory is a different story: it stays at the initial 8 GiB throughout the solving phase (that phase is dominated by model calls, memory consumption is small), until the evaluator needs it and doubles twice to 32 GiB. When the task completes the environment is torn down, and quota immediately drops to zero, without waiting for conservative downscale.

The value of this example is that it proves the necessity of elastic quotas: "the same environment holds widely different quotas at different moments, which no fixed allocation could match."

4.6 Overall Score: Init Share 53.4% → 9.1%

(§9.2) Same workload run on four backends:

MetricE2BDockerAgentENVWeEnv
Env init share of iteration time53.4%39.5%—9.1%
Task exec share of iteration time20.2%44.0%—74.0%
Median initialization150.6 s111.0 s59.7 s10.6 s

The three baselines' initialization shares range from 28.9%–53.4%; WeEnv compresses it to 9.1%, with task execution rising to 74.0%—the original summarizes: "the iteration finally spends its time on solving the tasks rather than on preparing the environments."

Median initialization 10.6 s vs 59.7 / 111.0 / 150.6 s, i.e., 5.6–14.2×. The distribution shape is also worth noting: Docker's curve is bimodal—38% of initializations hit a warm container on the node and complete within 30 s, the rest pay 100–300 s for full cold creation; while WeEnv "keeps the entire distribution low."

Robustness experiments (this is where the paper's quality shows; sensitivity was tested in three dimensions):

VariationWeEnvE2BDockerAgentENV
Switch harness to Codex, P50 task time108 s260 s228 s178 s
Same, median initialization13.7 s160.2 s126.8 s87.3 s
Qwen3-14B + 8 servers, P50 task time188 s288 s253 s207 s
Same, median initialization13.9 s150.7 s108.2 s60.4 s

When switching datasets (Rust / C++ / Go / JS), WeEnv is also fastest: 66.9 / 131.5 / 78.9 / 105.3 s. The original's explanation of the differences is apt: WeEnv's savings come mainly from the environment side, roughly constant per task, so the relatively faster the task itself executes, the greater the relative gain (up to 3.4× on Rust, Go), while C++ is compilation-intensive and task time itself dominates, narrowing the advantage over the strongest baseline to 1.2×—but still ahead.

4.7 WeEnv's Two Technical Criticisms of AgentENV

WeEnv explicitly compares AgentENV's storage design in §9.2; both criticisms point to architecture rather than implementation:

(1) Block-level conversion is not cache-friendly.

"it converts file-level layers into block-level ext4 layers with overlaybd, and the conversion is not cache-friendly: the same layer may map to different blocks across images, so the converted blocks cannot be shared by digest, an issue also observed in file-oriented image services [31]."

Compare its own approach: "WeEnv instead caches at the layer level, where a cached layer is reused wherever its digest matches, regardless of the enclosing image."

This criticism is structural: file-level layers have content-addressed digests, naturally globally deduplicable and shareable; once converted to block-level ext4, "the same upper-layer content" may land on different blocks in different images, breaking cross-image digest-based reuse. This is an inherent cost of the "block-level route," not an AgentENV implementation bug.

(2) Narrow optimization scope.

"AgentENV optimizes solely the initialization, whereas WeEnv spans the full lifecycle, rethinking packaging and resource provisioning as well."

Against the "three narrow gates" diagram in Section I: this is exactly the inevitable result of AgentENV betting on "startup + state"—it puts almost all engineering effort into initialization and snapshot/fork; packaging and provisioning are outside its scope.

4.8 WeEnv's Costs and Boundaries


V. Moonshot / AgentENV: Turning "an Existing Image" into "a Snapshotable, Forkable, Pausable Environment"

5.1 Positioning: The Only Open-Source Product Among the Three

AgentENV comes from Moonshot AI, code under GitHub org kvcache-ai. The repository states:

"AgentENV (AENV) is a platform for running agent environments at scale, powering agentic RL training for Kimi K3." (README:22)

Verifiable repository facts (local checkout):

ItemValueLocation
LicenseMITLICENSE:1 (Copyright (c) 2026 AgentENV)
First open-source commit2026-07-25 feat: Initial open-source release of AgentENVgit log --reverse
Total commits235 (as of local checkout)git rev-list --count HEAD
Production model claimKimi K3, external link to github.com/MoonshotAI/Kimi-K3 technical reportREADME:22, 28
Storage crate metadataname = "overlaybd", version = "0.1.0", no repository / authors / license fieldsstorage/overlaybd/Cargo.toml:1-4
Disclosure of DeepSeekSearching all *.md for deepseek / dsec zero hitsSee §1.2

Its positioning is fundamentally different from the other two: DSec and WeEnv are "systems in papers"; AgentENV is "something you can directly install" (install.sh / aenv CLI / E2B SDK compatible). So when evaluating it, "what is in the code" is more credible than "what the README says"—below I will strictly separate the two.

5.2 Architecture: One Storage Subsystem Carries the Core, Plus Node Orchestration and a Go Control Plane

AgentENV's own architecture document states its focus bluntly (docs/src/internals/architecture.md:3):

"AgentENV runs AI agents inside isolated, snapshot-capable Firecracker microVMs. Its core is a storage subsystem that provides layered block devices mountable into VMs and ublk-backed memory snapshot restore. The system also includes a per-node orchestrator managing sandbox lifecycle, and a distributed control plane (gateway + scheduler) for multi-node routing."

Three-layer structure:

AgentENV Node
├── API (Axum) ── E2B compatible + reverse proxy
├── Orchestrator ── lifecycle state machine (Creating/Running/Pausing/Paused/Resuming/Killing)
├── Firecracker VM
│     ├── /dev/vda rootfs ── ublk /dev/ublkbN ── overlaybd (upper writable + layer 0..N read-only)
│     ├── /dev/vdb extra drive ── ublk ── overlaybd
│     └── VM memory ── ublk read-only memory block device ── overlaybd mem layers
│            (multiple sandboxes of the same snapshot share this one device via reference counting)
└── Cluster side: Gateway(:8080) ── gRPC ── Scheduler(:9090)

Note the shape of the memory path: the memory snapshot is not a file, but another ublk block device, using the same overlaybd layering logic as rootfs. This is where AgentENV differs most from the other two—it turns "memory" into a block device that can be layered, read on demand, and shared by multiple sandboxes.

5.3 Mechanism One: Block-Level Layered Storage—LSMT Format + ublk Devices

overlaybd's format (architecture.md:47-67): not a simple layered tar, but an LSMT (Log Structured Merge Tree) format:

ublk device lifecycle (architecture.md:69-85):

  1. UVMUblkCtrlBuilder sends ADD to /dev/ublk-control via io_uring's UringCmd;
  2. Kernel allocates device ID, creates /dev/ublkcN (control) and /dev/ublkbN (block);
  3. One worker thread per queue, each with thread-local AsyncIoRing and slab-allocated I/O slots;
  4. Kernel delivers block I/O to the mmap'd ublksrv_io_desc array, processed asynchronously in user space;
  5. I/O buffers use AutoRegBuffer (sparse buffer table for zero-copy, kernel 6.8+) or traditional UserBuffer.

A detail only known from hitting the pit (architecture.md:92): ublk devices are uniformly held by a resident daemon uvm-ublk-daemon, with node side sending RPC over Unix domain socket. The reason is very practical:

"All remote image I/O (OSS/registry reads) is dispatched from the per-queue runtimes onto a dedicated per-ImageService remote-io runtime via RuntimeDispatchFile, because opendal/reqwest pin pooled HTTP connection tasks to whatever runtime polls them and per-device runtimes are dropped on device teardown."

That is: tokio's HTTP connection pool pins tasks to "the runtime that first polled them," and per-device runtimes are dropped on device teardown—if remote I/O ran directly on per-queue runtimes, when the device is destroyed, tasks in the connection pool become orphans. So all remote I/O is dispatched to an independent, resident remote-io runtime (controlled by ublk.overlaybd.remote_io_workers, default 4, see config/default.toml:313). This is a pure runtime ownership problem, not an algorithm problem, and it is exactly the dividing line between "it runs" and "it explodes after running for a while."

5.4 Mechanism Two: Snapshot, Fork, Resume—AgentENV's True Differentiation

This is the deepest part of AgentENV among the three, and where it puts all its engineering effort.

(a) pause is just a "seal layer + reopen upper" action (architecture.md:65)

"ImageFile::create_snapshot_and_restack() is the primary pause path. It seals the live upper layer via LSMTFile::close_seal_and_reopen() so the upper becomes the newest lower layer, then reopens a fresh writable upper in place."

Note the consistency of this design: after sealing, the original upper becomes the youngest read-only layer, and a new writable upper opens in place—so after pause the environment can immediately continue running, while the entire filesystem increment it just accumulated has become a reusable lower layer. Snapshot is not an "export" action, but a rearrangement of layer stack pointers.

(b) What a snapshot retains (docs/src/concepts/snapshots/index.md:13-24)

(c) fork is a first-class API, not a wrapper around "copy snapshot"

The route actually exists in the OpenAPI contract and generated code: POST /sandboxes/{sandboxID}/fork (src/api/openapi.yml:2045, src/api/generated/src/server/mod.rs:73). The contract semantics are clear:

"Number of forked sandboxes to create. All forks boot from the same snapshot, so the snapshot is captured once regardless of count. Each fork succeeds or fails independently; the outcome of each is reported in its entry of the response list." (src/api/openapi.yml:836-838)

Several judgments worth recording at the implementation layer (src/orchestrator/service.rs:629-860, src/sandbox/firecracker/sandbox.rs:302-500):

ModeBehavior at forkLocation
exclusive (default)Each child creates an independent COW subvolume, and drive replacement relationships are registeredsrc/api/impls/sandbox.rs:221-243; enum default see src/volume.rs:36-41
roReuses the same volume id (read-only, multi-owner safe)src/api/impls/sandbox.rs:245-248

The underlying reservation semantics are consistent: exclusive is single-owner reservation, conflicts error directly (src/volume.rs:591-600); ro records multiple owners (:579-589). Documentation matches code (docs/src/concepts/volumes/sandbox-forks-and-snapshots.md:12-16).

(d) Memory snapshot bypasses userfaultfd, goes through block device—this is the highest technical content here

architecture.md:112-120 records a clear architectural turn:

"Memory snapshot restore uses ublk-backed overlaybd devices rather than userfaultfd. On resume, a read-only ublk device is created from the stacked memory overlaybd layers and passed to Firecracker as a BackendType::File memory backend. Firecracker mmaps the block device and COWs pages into anonymous memory on first write, so the underlying device is never modified."

Three consequences, each with practical value:

  1. Multiple sandboxes of the same template share one memory block device (reference counting): "Multiple sandboxes booting from the same snapshot template share a single memory ublk device via reference counting. This allows the Linux page cache to be reused across all sandboxes using the same memory image, significantly reducing I/O for concurrent launches from the same template."

→ This directly serves the scenario of "instantly bringing up a large batch of environments from one template," and sharing happens at the host page cache layer, requiring no deduplication algorithm.

  1. Copy-on-write does not copy the original device: the first write triggers Firecracker to COW the page into anonymous memory, the underlying read-only device is never rewritten—so one memory layer can safely serve N sandboxes simultaneously.
  2. Memory snapshots are incrementally layered: at pause, Firecracker generates a state-only diff snapshot; AgentENV queries Firecracker's dirty/present memory ranges, uses process_vm_readv to read selected memory, directly creates an OverlayBD memory layer; and stacks the parent layers of previous snapshots to form a complete layered memory image. → Sandboxes that repeatedly pause/resume pay only the increment each time.

But increments are not free. Layers accumulate continuously, so the code configures a layer budget: DEFAULT_MAX_OVERLAYBD_SNAPSHOT_LAYERS: usize = 32 (src/sandbox/firecracker/overlaybd_snapshot.rs:41, comment says it is the "base budget for runtime-owned snapshot lowers before merging"). Exceeding triggers compaction to collapse multiple layers.

This is the half most easily missed in "incremental snapshot" designs: DSec solves the same problem on the container side with offline layer collapsing (~3 GB threshold) (§3.5); AgentENV uses a 32-layer budget at runtime to contain it. Both independently discovered: a layering scheme must be paired with an answer to "what if there are too many layers," otherwise read amplification and mount overhead gradually eat the gains.

Corroborating the thoroughness of this turn: storage/uffd-core/ remains in the repository, but is excluded from the workspace build (architecture.md:120: "retained for reference but excluded from the workspace build"). That is, the userfaultfd path is abandoned, not optional.

(e) warm pool: moving two things off the resume critical path in advance

architecture.md:142:

"Snapshot resume can also use [pool.firecracker] to pre-spawn (network slot, Firecracker process) pairs. A warm entry transfers its network slot, process, and Firecracker CWD to the resumed sandbox, which avoids the spawn and API-socket wait in the resume critical path."

Config defaults are on (config/default.toml:338-344): pool.firecracker.enabled = true, startup_prewarm = true, fill_concurrency = 4. There is also [pool.block] reusing the OverlayBD ublk device pool (default.toml:331-336), and shared watermarks low_watermark = 2 / high_watermark = 64 (default.toml:321-325).

This is a very typical piece of "startup latency engineering": latency is not saved in the algorithm, but work is moved to idle periods (pre-spawn processes, pre-warm network slots, pre-create block devices) and then ownership is transferred on the critical path.

5.5 Mechanism Three: Kernel-Side Density Measures—DAMON + Page Reporting, and Enabled by Default

AgentENV's guest kernel command line (config/default.toml:33) contains this string:

page_reporting.page_reporting_order=6
damon_reclaim.enabled=Y
damon_reclaim.min_age=100000
damon_reclaim.quota_ms=20
damon_reclaim.quota_sz=1073741824
damon_reclaim.quota_reset_interval_ms=500
damon_reclaim.wmarks_high=990
damon_reclaim.wmarks_mid=990
damon_reclaim.wmarks_low=200
damon_reclaim.skip_anon=Y
damon_reclaim.wmarks_interval=1000000

And balloon-side config—src/sandbox/firecracker/instance.rs:417-434 explicitly enables free page reporting:

pub async fn set_balloon(&self) -> Result<()> {
    let balloon = Balloon {
        ...
        free_page_hinting: None,
        free_page_reporting: Some(true),
    };

The config file comments also specifically explain the purpose of this combination (config/default.toml:31-33: "Editing it replaces the shipped boot and DAMON reclaim settings"), while src/sandbox/firecracker/config.rs:38-39 explains the meaning of each wmarks item.

Here appears a technical echo that must be pointed out: DSec §5.2 uses the same set of kernel mechanisms—"We therefore combine DAMON (Data Access MONitor) with virtio-balloon free-page reporting", and empirically proves DAMON + balloon FPR reduces time-integrated memory consumption by 21.2%.

The two take different values on the same parameter, which is interesting:

DSec §5.2AgentENV default.toml:33
Page reporting granularity"operates on order-9 pages, corresponding to 2 MiB regions with 4 KiB base pages" (default, adjustable)page_reporting.page_reporting_order=6 (256 KiB)
DAMON cold page reclaimEnabled (with balloon FPR)damon_reclaim.enabled=Y, skip_anon=Y

Convert: order-9 = 2⁹ × 4 KiB = 2 MiB, order-6 = 2⁶ × 4 KiB = 256 KiB—AgentENV makes reporting granularity 8× finer. For microVM scenarios this is reasonable: finer granularity means more free pages can be assembled for reporting and reclamation is more precise, but scan overhead is also greater. This is a typical engineering choice "papers won't write, only known after tuning."

It must be clear: identical kernel parameters alone cannot prove code homology (DAMON + FPR is a combination any team doing high-density Firecracker deployment would encounter). But together with the storage component relationship already disclosed in the DSec paper in §1.2, it forms a coherent picture: these two use the same technical vocabulary and the same batch of kernel features at the two "kernel boundary layers" of block storage and memory reclaim. Their platform layers (scheduling, lifecycle, security) remain completely different.

5.6 Mechanism Four: E2B Compatibility + an Honest Control Plane

External interface directly compatible with E2B:

POST   /sandboxes              create
GET    /sandboxes              list
GET    /sandboxes/{id}         metadata
DELETE /sandboxes/{id}         delete
POST   /sandboxes/{id}/pause   pause (snapshot)
POST   /sandboxes/{id}/resume  resume from snapshot
POST   /sandboxes/{id}/fork    fork (N children)
GET    /nodes, /nodes/{id}     node view
ANY    /proxy, /proxy/{path}   HTTP/SSE/WebSocket reverse proxy into sandbox

(docs/src/internals/architecture.md:181-192; fork is not in that list, it is defined by the contract: src/api/openapi.yml:2045)

How "E2B compatible" is specifically achieved: src/api/openapi.yml never mentions "e2b"—compatibility is not through an E2B-specific spec, but through aligning route shapes, request headers, and ID formats:

Compatibility pointLocation
Route header aliases e2b-sandbox-id / e2b-sandbox-portsrc/api/proxy.rs:87-92 (parsing at :818, :825, stripping before forwarding at :973, :975)
Traffic token header e2b-traffic-access-tokensrc/api/impls/auth.rs:14
API key prefix e2b_src/api_key.rs:20
MMDS field names / template step error shapes / default ready command aligned with E2Bsrc/sandbox/firecracker/mmds.rs:68, src/template/step_executor.rs:473, src/template/runner.rs:28

And it clearly states what is not compatible (docs/src/integration/e2b.md:96-97): "AgentENV supports the TypeScript SDK's volume create, list, and delete operations... The E2B SDK's direct volume content API is not supported."—A compatibility layer proactively marking its gaps is far more useful than claiming "fully compatible."

There is also [custom_extension] lifecycle hooks (default.toml:131-136), [sandbox_proxy].domains hostname-based routing, and the aenv CLI (Auth / Pull / Build / Start / Exec / Connect (alias cn) / Pause / Resume / List (ls) / Delete (rm) / Timeout / Snapshot (snap) / Template / Volume, see crates/aenv/src/main.rs:20-65).

The control plane is an independent Go module (services/):

Client ──HTTP──> Gateway(:8080) ──gRPC──> Scheduler(:9090)
                     │
                     └──proxy HTTP──> Node A / Node B (:8000)

Regarding "where state lives," there is a common misreading that must be corrected for AgentENV. architecture.md:223 says:

"Limitations: All bindings are in-memory (lost on scheduler restart). After a scheduler restart, bindings are rebuilt from new sandbox creations plus the next heartbeat roster from each runtime node... but binding persistence is still not replicated."

This sentence is too broad; copying it verbatim underestimates this control plane. The actual situation is:

FactLocation
DefaultBindings in memory (NewInMemoryBindingStore)services/scheduler/cmd/main.go:165
But optional Redis backendNewRedisBindingStore, Lua scripts for atomic writesservices/scheduler/internal/redis_store.go:28, 103-113
And documented HA shapeOne primary (with scheduler.redis_addr) + multiple --query-only replicas, the latter serving only LookupNode from Redisservices/README.md:293
Still a real gap"Artifact store state is still in-memory and is not covered by this HA mode."services/README.md:293

So the accurate statement is: AgentENV's control plane has an optional Redis binding store and query-only replica scheme, but artifact indexing remains only in memory under HA mode; the sentence in architecture.md:223 is too broad or even outdated.

Even so, this control plane's positioning is clear: AgentENV's focus is on in-node storage and snapshots; the control plane only handles "routing requests to the correct node." This looks similar to DSec's repeatedly emphasized design orientation that "placement engine and watcher need no persistent state," but the paths are opposite: DSec is actively designed to be stateless and validates with periodic IaC rebuilds; AgentENV is stateful by default first, then adds an optional Redis backend, and explicitly admits HA does not cover artifact indexing.

5.7 Verifiable vs. Merely Claimed: Singling Out README Numbers

AgentENV has no accompanying paper, so its performance claims can only come from the README. I distinguished "code-supported" from "text-only" item by item:

ClaimSourceCode-level evidenceVerdict
boot/resume < 50 ms, pause < 100 msREADME:29 (same sentence in docs/src/getting-started/overview.md:9)Has benchmark suite (crates/benchmarks/benches/snapshot_benchmark.rs: snapshot_resume, snapshot_resume_cold, concurrent_resume, concurrency CONCURRENCY = 50), but only times, no thresholds, and no result files in repo; mechanism side has warm pool (src/sandbox/firecracker/pool.rs comment: "entry to skip process spawn and API socket polling on the critical path") + pool.firecracker.startup_prewarm = true⚠️ Claimed value, no threshold/no results
9.6× memory overcommit ratio (production)README:31Mechanisms complete (balloon FPR + DAMON + memory device reference-count sharing + control plane allows >100% allocation: services/shared/config/config.go:46-48 max_memory_allocated_percent "can exceed 100 (overcommit)"); but the string 9.6 appears elsewhere in this repo only in unrelated crate version numbers in Cargo.lock, with no measurement or config item corresponding to it⚠️ Specific value unsupported; and docs/src/getting-started/overview.md:11 deletes "9.6x" in the same paragraph
snapshot < 100 ms, even with heavy disk modificationREADME:30Incremental capture mechanism real (src/sandbox/firecracker/sandbox.rs:1345-1351: state-only diff snapshot + get_dirty_memory_ranges()); heavy-disk scenario has bench variant (bench_snapshot_creation_1gdisk)⚠️ No 100 ms assertion
1.5 million images in productionREADME:28External link MoonshotAI/Kimi-K3 technical report, not in this repo⚠️ Cannot verify locally
native snapshot / forkREADME:30✅ Complete implementation: pause restack, POST /fork, per-child independent results, volume COW subvolumes, graded failure semantics, dedicated tests✅ Verifiable
on-demand loading of OCI imagesREADME:28✅ image.resolver.convert_standard_oci = true (default.toml:76), registryfs_v2 backend, image.cache.remote_blocks.max_size_gb = 100 (default.toml:95-97)✅ Verifiable
local disk as bounded cache, hot retained, cold evictedREADME:28✅ image.cache.capacity_gb = 100 + GC (interval 1800 s, min free 600 s, high watermark 0.95 / low watermark 0.70, default.toml:92-111)✅ Verifiable
P2P accelerationdefault.toml:113-119⚠️ Disabled by default (enabled = false), and comment self-describes "EXPERIMENTAL: P2P has not been tested in production. Keep it disabled in production unless the deployment accepts that operational risk."; another factory config self-contradiction: [snapshot].p2p_enabled = true (:228) while [p2p].enabled = false (:117), the former does not take effect due to forced downgrade in src/p2p/config.rs:36-40⚠️ Self-declared unverified

The two most noteworthy points in this table:

  1. 9.6x only appears in the README, and is deleted in the corresponding docs paragraph. (README:31 vs docs/src/getting-started/overview.md:11: the latter retains "Memory ballooning returns reclaimable guest memory to the host," but rewrites "achieving a 9.6x memory overcommit ratio in production" to "sustaining high overcommit.") This inconsistency is an observable difference within the documentation—I cannot determine whether it is deliberate or an oversight, but citing this number should note that the source is README only.
  2. P2P is disabled by default and self-declared not production-verified—an open-source project's README treats P2P as one of its selling points, while the config file explicitly says not to enable it in production. Such self-limitation is a good signal (shows maintainer honesty), but it also means P2P cannot be counted as an AgentENV capability during evaluation.

5.8 AgentENV's Costs and Boundaries


VI. Horizontal Comparison: Three Institutions' Three Answer Sheets

6.1 Summary Table

DimensionDSec (DeepSeek + Tsinghua)WeEnv (Tencent WeChat)AgentENV (Moonshot)
Material formTechnical report (31 pages)Paper (16 pages)Open-source codebase
Isolation baseFnCall / container / Firecracker microVM / QEMU full VM (4 kinds)Container (only)Firecracker microVM (only)
Container wrapped in VM?Yes (containers run inside QEMU/libvirt VMs)NoN/A
Environment packaging unitbase + workspace + toolkit (EROFS layers)layer group + environment planDoes not do packaging
Who builds environmentsAgents themselves online (pack_diff)Offline cut by change rate + repository-granularity packagingExternal (BuildKit builds templates)
On-demand loadingContainers: EROFS multi-device; microVMs: overlaybd + ublk, 256 KiB chunk + local second-level cacheFUSE + custom layer format (tail TOC + contiguous payload)overlaybd (LSMT) + ublk, remote blocks cache
Content addressing granularityFile-level (EROFS) / block-level (overlaybd)File-level (layer digest reusable across all images)Block-level (LSMT segment)
Density measuresvirtio-pmem+DAX, DAMON+FPR, CPU QoS (core scheduling)Elastic provisioning (200 ms cgroup sampling)balloon FPR + DAMON + memory block device reference-count sharing
Snapshot / forkpause/resume (container docker pause; microVM snapshot+kill process)Not a core mechanismFirst-class: restack pause, POST /fork, memory device sharing
Coordination with trainerDeep co-design (rollout moved out of preemptible GPU pool)Robustness only (clean evaluation environment prevents contamination)None (positioned as general platform)
Adversary modelMost complete (AppArmor + eBPF network allowlist + case postmortems)Architectural (independent clean evaluation environment)Not public
External interfacelibdsec (Python SDK)Integrated with SlimeE2B-compatible HTTP API + aenv CLI
Scale evidence160 nodes/scale unit, 3M sandboxes/day, 380K concurrency, >5,000 creations/s"Deployed at WeChat," scale undisclosed; evaluation 4–6 H20README claims 1.5M images; only third-party measurement is weakened
Performance measurement credibilitySelf-measured, with ablations, 10-node experimental clusterSelf-measured, with 3-dimension sensitivity analysis, and actually measured three competitorsNo independent measurement

6.2 They Are Not on the Same Track: Three Orthogonal Bets

Reading the summary table vertically reveals something easily overlooked: these three have hardly ever clashed head-on on the same technical decision.

The orthogonality of the three bets shows in several concrete places:

QuestionDSec's answerWeEnv's answerAgentENV's answer
Where does the environment come from?Agents build onlineOffline cut by change rateDoesn't care
How does the environment reach the node?EROFS/overlaybd on-demand from 3FSCustom format on-demand from registryoverlaybd on-demand from registry/OSS
After it's running?Coordinate pause/resume with trainerElastically adjust CPU/memory quotaSnapshot, fork, share memory pages
Who is most dangerous?The model itself (reward hacking)Contaminated evaluation signalsNot stated

So "who is better" is the wrong question. The right question is: which cut you need to make, and whose range covers it.

A coarse-grained selection guide:

If your constraint is…Closer answer
Security-sensitive tasks / full OS environments / Android / graphics / GPU kernelDSec (four-backend spectrum, no alternative)
SWE-type container-reachable workload, pain is "iteration time eaten by environments"WeEnv (full lifecycle + elastic quota)
Need environment state portability, fork branching, or an E2B-compatible self-hosted baseAgentENV
Need preemption coordination with RL trainer (GPU jobs frequently preempted)DSec (the only one that did this)

6.3 The Deepest Disagreement: The Three Draw "Environment" Boundaries Completely Differently

This is, I think, the most memorable point of the entire article. All three say "environment," but their denotations of the word "environment" are fundamentally different:

What "environment" isImplied responsibility boundary
DSecA programmable artifact (base + workspace + toolkit, agent can pack_diff anytime)Environment is something produced at runtime; the platform must support online building
WeEnvA software supply chain requiring governance (4 layer groups with different change rates)Environment is a release artifact requiring continuous maintenance; the platform must manage its full lifecycle
AgentENVAn input (an existing OCI image)Environment is a given premise; the platform only turns it into a fast sandbox

This definitional difference directly determines whether WeEnv's criticism holds. WeEnv says AgentENV "optimizes solely the initialization"—from AgentENV's own boundary definition, this is not a defect, but its definition itself: packaging is not its responsibility. Conversely, DSec's pack_diff is possible precisely because it does not hand packaging to an external pipeline.

One-sentence summary of the three's division of labor:

DSec manages "how environments are created"; WeEnv manages "how environments are maintained and scheduled"; AgentENV manages "how environments are started and copied."

The three can be stacked, not substituted—which also explains why DSec would open-source its storage layer into AgentENV's repository: at the "block device path" layer, the two share; above it, each builds its own tower.


VII. The Only Cross-Institution Measurement: WeEnv's Bench, and Four Qualifications That Must Be Added

In the entire article, only one set of data puts products from three different institutions on the same experimental bench—WeEnv §9. It is the only public source for understanding AgentENV's relative performance. Therefore its qualifications must be itemized, otherwise it is easily misused.

7.1 The Measured Data Itself

Metric (SWE-Smith-Python, Qwen3-8B, Claude Code)E2BDockerAgentENVWeEnv
Median environment initialization150.6 s111.0 s59.7 s10.6 s
Median initialization after switching to Codex harness160.2 s126.8 s87.3 s13.7 s
P50 task time after switching to Codex260 s228 s178 s108 s
Qwen3-14B + 8 servers, P50 task time288 s253 s207 s188 s

AgentENV's position in this table is clear: fastest among all baselines, but behind WeEnv. This itself is favorable to AgentENV—it is the only one of the three baselines below 60 s on initialization (the other two are 111 s and 150.6 s).

7.2 Qualification One: AgentENV's Core Mechanism Was Disabled in the Experiment

This sentence must be quoted verbatim (WeEnv §9.1):

"Since our servers lack ublk kernel support, the on-demand fetching is disabled in our experiments."

ublk is the physical carrier of AgentENV's on-demand loading (the entire path in §5.3 and §5.4 is built on ublk block devices). Servers lacking ublk kernel support means AgentENV's block-level on-demand pull completely did not run—and this is precisely its core advantage over "pulling full images."

Conclusion: The 59.7 s number is not AgentENV architecture's performance, but "what AgentENV degrades to on a machine without ublk."

7.3 Qualification Two: WeEnv Also Modified AgentENV for Speed

Immediately following in the same paragraph:

"We tune AgentENV for better performance. The file-level layers are converted into block-level ext4 layers offline and pushed to the registry, so that initialization fetches the converted layers directly and bypasses the conversion during initialization, reducing its initialization latency by 25%."

That is: moving conversion work off the initialization path, buying 25% initialization speedup.

This point is often overlooked, but it is important because it shows WeEnv's operation is not simply weakening the opponent: they both disabled AgentENV's on-demand loading (−) and moved its initialization-path conversion overhead away (+). The net bias direction is unknown—we cannot calculate from the paper whether these two offset to optimism or pessimism.

The honest conclusion is: The 5.6–14.2× multiple holds under this specific setup of "no ublk + offline pre-conversion." It is a valid measurement, but it can neither be extrapolated to "AgentENV architecture is 5.6–14.2× slower," nor used to infer DSec's performance.

7.4 Qualification Three: DSec Is Completely Absent from This Bench

WeEnv's four backends are E2B, Docker, AgentENV, WeEnv. DSec is not among them—the two papers are 6 days apart (09-19 vs 09-25), and I checked both reference lists item by item; neither cites the other.

So:

7.5 Qualification Four: AgentENV's "<50 ms" and the Measured "59.7 s" Are Not Directly Contradictory, but Cannot Be Aligned

AgentENV README:29 claims "Snapshot-backed environments boot or resume in under 50 ms and pause in under 100 ms"; WeEnv measured median initialization at 59.7 s.

These two numbers differ by three orders of magnitude, seemingly contradictory, but likely measure two different things:

The real problem is not contradiction, but unalignability: AgentENV never defines the start and end points of "boot," which path it measures, under what cache state, and leaves no benchmark scripts in the repository. So these two numbers cannot be placed on the same coordinate axis. This is not unique to AgentENV—as long as a project's performance claims are not reproducible, it cannot participate in cross-institution comparison.

7.6 One-Sentence Conclusion

WeEnv's evaluation is currently the only cross-institution data; its value is giving a lower bound under restricted conditions. But when using it, four labels must be attached: AgentENV on-demand loading disabled, AgentENV offline pre-conversion speedup 25%, no DSec in the experiment, the two compared metrics have different definitions.


VIII. What Each of the Three Institutions Left That Can Be Directly Taken Away

Setting aside "who wins," there are six conclusions in these three materials that can be directly transferred to any team doing agent infrastructure. Each is labeled with its source institution—this is also the practical use of the "three institutions" framework: they are empirical experiences from three different situations.

8.1 Measure the Environment Tax First, Then Decide Which Cut to Make (Source: WeEnv)

WeEnv's methodology can be copied directly: split RL iteration into Train / Env init / Task exec / Others four segments, normalize by total latency, then split rollout into initialization and task execution. After doing this, you know whether to invest in training, environment, or task.

WeEnv's measured results (53.4% / 39.5% initialization share, training only 15.4%–24.0%) will likely differ for other teams, but "split and measure first" itself is extremely low-cost and high-information.

8.2 Sparsity Is a Dependable Assumption, but Granularity Depends on the Magnitude You Measure (Source: WeEnv + DSec)

Two independent instruments give same-direction but different-magnitude numbers:

Measurement objectAccess share
WeEnv (file-level layers, SWE tasks)Median 0.81%, p99 6.02%
DSec (container image overall, by language)4.2% – 13.3%

Both are correct because the measured granularity differs: WeEnv counts "artifact bytes actually read by the task," DSec counts "fraction of image data accessed at runtime."

Design implication: Do not copy numbers; copy the order of "measure first, then set granularity." And this magnitude difference exactly explains why both routes work—file-level layers are naturally coarser than block-level, so file-level schemes' gains depend more on "how well layers are cut," while block-level schemes' gains depend more on "how fast on-demand reads are."

8.3 Treat "Change Rate" as the First Principle of Packaging (Source: WeEnv + DSec)

Two institutions independently reached the same packaging principle:

Two actionable corollaries:

  1. One code repository per artifact, putting per-task differences (checkout a commit, apply a faulty patch) at initialization is often the highest cost-performance cut—WeEnv measured this mutation at median 0.43 s, while installing harness takes 20 s (46× difference).
  2. Changing packaging granularity may require only a small change (DSec's 30 lines of Go), but the gain is order-of-magnitude. First measure "which components change frequently," then decide where to cut.

8.4 Failures Must Be Explicit, Never Silently Degraded (Source: WeEnv + AgentENV)

WeEnv lists it as the third scheduling principle, with a reason weightier than any performance argument:

"A trajectory silently corrupted by starvation is worse than a lost one, as it feeds wrong signals into training."

So when OOM occurs with no tier left to raise, it chooses to terminate the task and report the cause, letting the RL framework resample, rather than letting it run sick to completion.

AgentENV gives the engineering form of the same principle at the implementation layer: failures must be graded—

Failure typeHandling
Recoverable (capture/fork error)Restore source sandbox to Running, source unharmed
TerminalOnly then require tearing down source sandbox
Child sandbox start failureMust not drag down already-successful siblings (fork_sandbox_keeps_successful_siblings_when_one_start_fails)
Volume subvolume creation failurecleanup_fork_volume_children() rollback

These two things are the same principle: in the face of training data correctness, "partial success" is far more dangerous than "explicit failure."

8.5 Treat Agents as Adversaries, Not Users (Source: DSec + WeEnv)

DSec §6.4's case list should be read by everyone doing agent infrastructure. The most valuable part is not its mitigations (AppArmor + eBPF), but what the failures look like:

WeEnv gives an architectural answer from the other side: evaluation must run in an independent, clean environment, "so that rollout-side modifications cannot contaminate the reward computation, guarding against reward hacking."

DSec also adds an easily overlooked governance measure: builder and runtime use different accounts, and build-time residual data is cleared from the writable layer before packaging, avoiding bringing reference answers into the image.

Actionable corollary: Treat "from which unintended channels answers may leak" as an attack surface requiring continuous enumeration (internal sockets, logs, overwritten interpreters, package proxies, process filesystem), not one-time data cleaning. Two independent institutions both confirm this is a continuous adversarial process—WeEnv even writes it as a packaging cost item ("evaluator must be frequently updated against reward hacking; each change touches affected images, possibly taking days").

8.6 Startup Latency Relies on "Doing in Advance + Ownership Transfer," Not Algorithms (Source: AgentENV)

AgentENV's warm pool is a very clean paradigm (architecture.md:142, default.toml:338-344):

This and DSec's placement design are two faces of the same philosophy: DSec uses "local view overlaying its recent placements not yet reflected in periodic watcher snapshots" to eliminate coordination latency; AgentENV uses warm pool to eliminate creation latency. Neither is "making a step faster," but "moving a step off the critical path, or making it require no coordination."

8.7 On Whether the Control Plane Should Hold State (Source: DSec + AgentENV Contrast)

The two take opposite approaches, but both write their reasons clearly, which is more important than choosing a side:

DSecAgentENV
ApproachExplicitly designed to require no persistent stateDefault in-memory, but provides optional Redis binding store + query-only replicas (services/README.md:293, services/scheduler/internal/redis_store.go:28)
ReasonMake apiserver / placement / watcher instances freely addable, removable, replaceable, with no expensive recovery steps; validate with periodic IaC rebuildsLeave state authority to nodes; control plane only routes; HA needs filled by optional backend
CostRequires rebuild capability and validation mechanism (BGP + multiple instances + periodic cluster reset)architecture.md:223's statement is too broad; the real gap is artifact indexing still only in memory under HA mode

Common point: Neither treats the control plane as the sole authority for state. The difference is opposite paths—DSec is actively designed stateless and proves rebuild capability with periodic IaC rebuilds; AgentENV is stateful by default, then adds an optional Redis backend, and explicitly admits HA does not cover artifact indexing. The former trades constraints for elasticity; the latter trades options for elasticity.


IX. Closing

Part 2 of 3 · background

How Xiaomi MiMo-V2.6 Actually Does RL

One release read down to its config values. This one takes three releases and reads them against each other: what DeepSeek, Tencent WeChat and Moonshot each built for the same bottleneck, where their designs genuinely diverge, and which of their numbers you can actually check. The environment tax that all three measure is the same one Part 1 named and Part 2 traced through a single stack.