Three institutions, three released systems, one shared diagnosis: the environment, not the agent, is where the time goes. A side-by-side teardown of DeepSeek's DSec, Tencent WeChat's WeEnv and Moonshot's AgentENV.
Project Institution Material Form Time DSec (DeepSeek Elastic Compute) DeepSeek-AI + Tsinghua University Technical report, arXiv:2609.22978v1 [cs.DC], 31 pages 2026-09-19 WeEnv Tencent WeChat AI (WeChat AI, Tencent) Paper, arXiv:2609.30766v1 [cs.DC], 16 pages 2026-09-25 AgentENV (AENV) Moonshot AI Open-source codebase, GitHub org kvcache-aiFirst open-source commit 2026-07-25
The problem these three projects aim to solve has its best name in WeEnv: environment tax.
The WeChat team used their production-grade framework Slime to break down iteration time (WeEnv §1, §3.1, Figure 2):
| Backend | Train | Env init | Task exec | Others |
|---|---|---|---|---|
| E2B (microVM) | 24.0% | 53.4% | 20.2% | 2.5% |
| Docker (container) | 15.4% | 39.5% | 44.0% | 1.1% |
Original wording:
"initialization alone accounts for 53.4% of the iteration time with virtual machines (E2B) and 39.5% with containers (Docker), 2.6× the task execution itself with E2B and roughly on par with it under Docker, while training takes merely 15.4%–24.0%."
On E2B, environment initialization eats 53.4% of iteration time—2.6× task execution and more than 2× training itself. Training accounts for only 24.0%.
Another perspective: rollout (= initialization + task execution) accounts for 73.6% on E2B and 83.5% on Docker. In other words, the vast majority of this pipeline's time is not spent learning, nor solving problems, but preparing the place to solve problems.
DSec did not do an "iteration time share" breakdown. It measured the shape of the workload (§4, data sampled from one week in early 2026):
Bursty creation. Number of sandboxes created per task (Fig. 2):
| Sandboxes created per task | p50 | p90 | p99 |
|---|---|---|---|
| Containers | 2,528 | 7,969 | 16,388 |
| microVMs | 352 | 1,835 | 4,044 |
"The largest production jobs can request up to 32K sandboxes. These requests arrive within a short window because the training or evaluation batch cannot use an instance until its environment is ready."
Extremely sparse access. Actual runtime access ratio of container images by language (Tab. 3):
| Image type | C++ | Go | Java | JavaScript | Python |
|---|---|---|---|---|---|
| Fraction of data accessed | 8.7% | 13.3% | 9.2% | 4.2% | 6.0% |
| Image size | 4.9 GB | 4.1 GB | 12.1 GB | 9.6 GB | 6.0 GB |
The vast majority of image data is never touched from beginning to end.
Low fanout. Image reuse within a single task (Fig. 8, over 1.5 million containers / 390,000 microVMs): container images p50 = 3, p90 = 28; microVM images p50 = 1, p90 = 3. DSec's conclusion is blunt: the working set is too diverse for single-node caching to absorb, so every burst necessarily triggers image pulls.
Sparse CPU + long-lived memory. Fig. 5: about 90% of container and microVM sandboxes use on average only within 5% of their requested CPU—a natural justification for oversubscription. Fig. 7: container lifetime p50 = 17.4 min, p99 = 231.5 min; microVM p50 = 15.5 min, p99 = 213.9 min. CPU is idle, but memory and state remain pinned.
Putting the two sets of measurements side by side reveals something very convincing—two companies, without knowing about each other, observed the same set of structural facts:
| Structural fact | DeepSeek (DSec §4) | WeChat (WeEnv §3) |
|---|---|---|
| Sparse access | Only 4.2%–13.3% of image data touched at runtime | Median reads only 0.81% of artifact, p99 only 6.02%, no task exceeds 7% |
| Bursty creation | Up to 32K sandboxes per task; container p99 = 16,388 | A batch is admitted in two waves; initialization is on the critical path |
| Long-tail CPU demand | ~90% of sandboxes average ≤5% of requested CPU | Peak CPU p50 = 0.99 cores → p99 = 68.3 cores, nearly 70× span |
| Long-tail memory demand | Memory reclamation is the density bottleneck | 92% of tasks peak < 0.2 GiB, but p99 jumps to 7.0 GiB |
| Dramatic demand fluctuations | CPU intermittent during tool calls, memory persistently resident | Single-task CPU demand surges 47× within one second |
This table is the most important argumentative foundation of the entire article. It shows that the "environment tax" is not a defect of any particular implementation, but a structural property of the agentic RL workload. Therefore, the disagreement among the three is not about "whether to solve it," but about "which cut to make first."
This is the skeleton of the entire article. Here is the conclusion first:
Common disease: environment tax
(bursty creation · sparse access · long-tail demand · state residency)
│
┌───────────────────┼───────────────────┐
▼ ▼ ▼
DeepSeek/DSec WeChat/WeEnv Moonshot/AgentENV
"Platform + density" "Full lifecycle" "Startup + state"
│ │ │
Bet: one platform Bet: manage the Bet: turn "an existing
manages four isolation entire chain from image" into "a running
backends, coordinates packaging to environment," and make
with RL trainer, provisioning it snapshotable,
handles cluster forkable, pausable
elasticity
│ │ │
Scarce resource: Scarce resource: Scarce resource:
node density & iteration time startup latency &
cluster elasticity (the tax itself) state portability
│ │ │
Evidence: 160 nodes/ Evidence: init share Evidence: README claims
3M sandboxes/day 53.4% → 9.1% <50ms / <100ms / 9.6×
The following three sections dissect each of the three narrow gates one by one.
DSec is a technical report (not a paper submission), authored by DeepSeek-AI and Tsinghua University. First author Jialiang Huang is a Tsinghua PhD student who completed the work during an internship at DeepSeek-AI; corresponding author is Liyue Zhang. It serves the sandbox workload for all RL training and evaluation from DeepSeek V3.2 to V4.1 (§6 opening).
Its self-positioning is very clear, and it deliberately avoids a single abstraction:
"This interface is intentionally not a full semantic abstraction over all backends. Function calls, containers, microVMs, and full VMs have different startup costs, isolation boundaries, filesystem semantics, and operating-system capabilities. libdsec provides a unified access path and a similar operational model, but the caller remains responsible for selecting a backend that matches the workload." (§2.1)
This is the only design among the three that "does not bet on a single isolation mechanism."
DSec §2.4 gives actual numbers for a scale unit:
| Metric | Value |
|---|---|
| Nodes | ~160 CPU nodes |
| Cores | 30K cores |
| DRAM | ~250 TB |
| Images/layers | Several PB |
| Daily sandbox instances | ~3 million |
| Peak concurrency | ~380,000 |
| Creation rate | >5,000 instances/sec |
And the density operating point (§4.3): production stably runs 3,200 containers or 800 microVMs per node (one-day sampled peak was 1,048 containers / 524 microVMs).
This magnitude explains the aggressiveness of all DSec design. When the creation rate is 5,000/s, "pulling the full image" is not slow—it is a structural error.
DSec §3's architecture is divided into three parts, each deliberately made stateless and replaceable:
Cluster-level services (control plane)
sandbox ID encodes the edge it belongs to, so any apiserver instance can forward directly—this is why it can scale horizontally.Node runtime
An easily overlooked but informative design choice: FnCall and containers both run inside QEMU/libvirt VMs, rather than directly on the host (§3.3).
"FnCall and containers run inside QEMU/libvirt VMs rather than directly on the host. The VM provides an isolated kernel and network stack and serves as an additional security boundary between untrusted containers and the bare metal."
That is, a VM wrapped around the container—sacrificing density for "untrusted containers do not directly touch bare metal." This "rather have one more layer" orientation reappears in §6.4's security story.
Problem (§4.2): A sandbox's content can be decomposed into three parts—base image (OS-level dependencies), workspace (task code repository + task-specific dependencies), toolkit (frequently updated tools, e.g., DeepSeek Harness).
Single-week production data (Tab. 2):
| Backend | base images | workspaces | snapshots | Aggregate size |
|---|---|---|---|---|
| Containers | 11,266 | 102,171 | – | 82.8 TB |
| microVMs | 2 | 53,590 | 4,889 | 50.9 TB |
There are also 103 toolkits, and 67.8% of sandboxes need a workspace or toolkit beyond the base image.
If they were welded into a single OCI image: maintaining M bases, N workspaces, K toolkits, upgrading m bases requires rebuilding O(m·N), upgrading k toolkits requires O(k·N). The original example is intuitive: upgrading Toolkit T1, every monolithic image containing it must be rebuilt, even if not a single byte of base or workspace has changed.
DSec's answer (§5.1): Treat the three as logical layers with independent lifecycles, dynamically stacking them with overlayfs at creation time. Base at the bottom, workspace inserted above as a read-only layer, toolkit on top.
The "dirtiest" and most interesting implementation detail: modifying dockerd.
"We modify the open-source Docker daemon (based on the Moby project) to dynamically insert EROFS-backed lower layers at container creation time. Specifically, we pass the path of a pre-mounted EROFS layer and insert it into the overlayfs stack before mounting, placing it as the topmost lower layer so it can override files in layers below. This change is minimal, requiring only 30 lines of Go code." (§7)
30 lines of Go exchange O(m·N)/O(k·N) → O(m)/O(k). This is the highest cost-performance engineering change in the entire article.
Why EROFS instead of ext4/XFS: EROFS is designed for read-only data, eliminating write-related accounting, with a more compact layout, supporting compression while preserving random access. Key comparison:
"Unlike a tar.gz archive, EROFS can read and decompress only the compressed blocks covering the requested data, so the complete image does not have to be transferred and unpacked before use."
Measured gains (§8.3): Same workspace + toolkit, tar.gz per-sandbox extraction vs EROFS direct mount:
| Scheme | End-to-end completion time | Disk write traffic | Peak write throughput |
|---|---|---|---|
| tar.gz extraction | 79 min | ~5.5× | ~3.4× |
| EROFS mount | 45 min (1.76× speedup) | 1× | 1× |
Note the original's explanation of "EROFS peak CPU is higher"—not because overhead is greater, but because more sandboxes enter the tool-calling phase concurrently earlier. This self-clarification is very honest.
The microVM side gets the same composition model (§5.1 end): base and toolkit are packaged as independent versioned EROFS images, exposed to the guest as read-only block devices; inside the guest, rootfs uses overlayfs, with the EROFS mount point as lower and a directory on the ext4 writable disk as upper.
DSec's key judgment is "do not build a separate image distribution layer" (§5.3):
"Existing on-demand image distribution systems often combine a container registry with peer-to-peer delivery to prevent the registry from becoming a bottleneck. We instead host images on 3FS, which already supports our production training workloads at scale. This choice reuses the existing storage infrastructure and avoids deploying a separate image-distribution layer."
But 3FS has a characteristic that must be accommodated: very high large sequential read/write throughput, very poor small random I/O. This asymmetry directly determines three design principles:
Container path uses three EROFS features:
It also does something to "reduce mount count": continuously collapsing layers within a size threshold (~3 GB) offline into a pair of metadata/data EROFS images, while preserving overlayfs whiteout semantics to correctly express file deletion.
The microVM path is different, and the reason is hard (§5.3):
"Docker's overlay2 driver cannot use an overlayfs-backed data directory. An alternative would be to export the host-mounted filesystem to the guest via virtio-fs, but our Firecracker backend does not support this interface."
So microVMs go block-level: read-only base/toolkit layers still use EROFS, while the writable ext4 disk goes through OverlayBD, exposed via ublk. The cost is stated openly: "Unlike the multi-device EROFS path, ext4 metadata remains embedded in the block image, so metadata reads can trigger remote I/O." Mitigation is 256 KiB chunk fetching + local second-level filesystem cache—after a chunk is evicted from page cache, it remains in the local cache and does not need to go back to 3FS.
Measured gains (§8.2): 10 nodes, 8,192 containers, real evaluation workload:
| Scheme | Completion time | Cumulative disk writes |
|---|---|---|
| Docker cold pull (eager) | >60 min | >1,600 GB/node |
| EROFS on-demand | ~35 min | ~700 GB (57% less) |
| Docker full local cache (upper bound) | ~35 min | ~600 GB |
On-demand loading matches the theoretical upper bound of "full local precaching", while being 1.71× faster than eager and writing 57% less. This result is more meaningful than "how much faster"—it shows the on-demand path introduces no additional penalty.
This section is the most "only known from production experience" part of DSec.
Memory (§5.2 / §8.4). microVMs have two independent sources of memory waste:
Two complementary answers:
The cost is also stated clearly: pmem requires the guest to allocate struct page for the entire range; with 4 KiB pages + 64-byte struct page, metadata consumes 1/64 of pmem capacity—"a 128 GB pmem device requires 2 GB of guest RAM for this metadata." And cold access may trigger synchronous page faults.
madvise(MADV_DONTNEED). But scattered file pages cannot form high-order blocks, so DAMON samples access bits to find cold file pages exceeding an age threshold and evicts them via the kernel reclaim path; after reclamation, buddy coalesces them into high-order blocks satisfying FPR. Measured time-integrated memory consumption reduced by 21.2%, with negligible CPU overhead.The two mechanisms are complementary: pmem eliminates duplication, DAMON+FPR reclaims idle pages. After combination, instantaneous peak CPU rose from 26.5% to 41.4%—DSec does not hide this and gives operational advice: "In CPU-constrained deployments, operators may prefer to enable FPR alone and retain virtio-blk."
CPU QoS (§5.2 / §8.5). The problem scenario is specific: some tasks have strict per-step latency budgets (e.g., board-game agents with fixed per-step time limits), and "lowering best-effort priority" is simply not enough—because BE and LS may run on the two SMT sibling threads of the same physical core, still competing for execution resources.
Two-layer strategy:
SCHED_IDLE, yielding CPU when LS is runnable;Measured (§8.5, LS per-step latency under 50% BE load):
| Configuration | Latency inflation |
|---|---|
| No protection baseline | 45.2% |
SCHED_IDLE only | At best improves 3.4% |
SCHED_IDLE + core scheduling | 17.3% |
"SCHED_IDLE alone improves at best only 3.4%" is a very valuable number—it proves that priority adjustment alone is nearly ineffective against SMT interference, and core scheduling must be touched.
The residual 17.3% is also explained: mainly from turbo frequency reduction under multi-core high load, memory bandwidth and shared LLC contention, which core scheduling cannot address; and it explicitly says memory bandwidth isolation is not done because it is "already tolerable."
If you remember only one thing about DSec, it should be this: it designs the sandbox platform and the RL trainer as one system. This is something the other two did not do.
(a) Let agents build their own environments (§6.1): The title is "Build environments of Agents, by Agents, for Agents". The core interface is called pack_diff:
"at any point, an agent can checkpoint a sandbox by taking an incremental disk snapshot, which can later be restored as a new sandbox. This checkpoint-and-restore interface turns an interactive session directly into a reusable environment, allowing environments to be built, validated, and consumed on the same infrastructure without a separate image-building pipeline."
Note the positional relationship here: both WeEnv and DSec are solving "environment packaging," but in opposite directions. WeEnv offline decomposes components into layer groups and recomposes them; DSec lets agents interactively freeze a running sandbox into an environment online. The former optimizes "combinatorial explosion of releases," the latter optimizes "there is no build pipeline at all."
Supporting governance measures are equally specific: constraints on agent build rules (limiting performance impact on shared infrastructure), an internal platform for quality inspection and exporting standard formats, builder and runtime using different accounts, and clearing build-time residual data from the writable layer before packaging to avoid bringing reference answers into the image.
(b) Splitting rollout out of the preemptible GPU pool (§6.2): The old version put the agent loop together with model serving and the RL framework in preemptible GPU training pods. When the GPU job is preempted, the agent loop is lost but the sandbox remains, so recovery can only rely on command log replay for reconciliation (reusing recorded results for completed operations, avoiding duplicate side effects of non-idempotent commands).
From DeepSeek-V4.1 onward, rollout execution was moved onto DSec, split into two components:
Both run outside the preemptible GPU pool. The result is that the rollout lifecycle is decoupled from the trainer lifecycle; the worker container + agent sandbox jointly hold the complete rollout state and serve as the single source of truth; preempted GPU jobs do not need command-log replay to rebuild execution after reconnecting.
(c) Pause means reclaim (§6.3): When a GPU job is preempted, the RL framework proactively sends pause to all relevant sandboxes; DSec reclaims memory but retains execution state; any subsequent request to a paused sandbox transparently resumes before execution. The two paths differ:
| Backend | pause | resume |
|---|---|---|
| Container | docker pause freezes the process tree → enable memory.swap.max → memory.reclaim proactively reclaims anonymous and file pages | MADV_WILLNEED async prefetch on process memory mappings → docker unpause |
| microVM | Snapshot memory and execution state → terminate the Firecracker process to release runtime memory | Start a new process → restore snapshot and continue execution |
Note that microVM pause really kills the process, retaining memory via snapshot to disk—this is the source of "pause consumes almost no memory," and is also the contrast point for AgentENV using the same trick later.
DSec §6.4 is the most "production-flavored" section of the entire piece, titled "Agent Misbehavior and System Failures". It classifies risks into two categories: the task appears to pass, but the answer comes from an unintended channel (undermining training/evaluation validity), and agent actions damage the environment (endangering this task or other tasks on the same machine).
The failure cases listed in the original are all very specific:
Finding answers inside the sandbox
/bin/bash to bypass checks or inject commands into subsequent shell sessions;XFS_IOC_SWAPEXT—an ioctl that swaps the data extent mappings of two files, aiming to make the contents of a protected file readable through another fd. This attempt corrupted XFS metadata and caused filesystem shutdown.Finding reference implementations outside the sandbox
Ordinary commands crashing infrastructure
grep from the root directory, traversing /proc, reading /proc/kpagecgroup, triggering a kernel bug and crashing the kernel;yes, while chronus caches stdout for asynchronous retrieval by the user, thus capturing tens of GB of stdout.Mitigations (§6.5) only admit to solving part of the problem; the title itself states "These controls address only part of the problem and do not provide a general defense against destructive behavior such as triggering kernel bugs."
network_rules={"npm": False, "pypi": True}.Strategic implication of this section: DSec's adversary model is adversarial. WeEnv also uses mechanisms (independent clean evaluation environments against reward hacking), but that is architectural isolation; AgentENV currently has no equivalent security narrative. This difference is expanded in Section V.
WeEnv comes from Tencent WeChat AI (WeChat AI, Tencent), authors Yang Yu, Jing Lei, Shaoxun Zeng (co-first authors, Shaoxun Zeng corresponding), paper arXiv:2609.30766v1, 16 pages. The conclusion is one sentence: "WeEnv is deployed for agentic RL at WeChat."—already online, but no scale numbers disclosed.
Among the three, WeEnv's stance is the most "management school":
"The root cause is the lack of a full-lifecycle solution to environment management." "WeEnv manages environments across packaging, initialization, and provisioning."
It also makes two choices the others did not:
WeEnv builds on container tooling: each environment is instantiated as a container, §5). No microVMs, no full VMs.Plus a methodological uniqueness: it is the only project that actually measured AgentENV as a baseline. This makes it the only source of cross-institution measured data among the three, and also brings fairness issues that must be flagged (see 4.7 and Section VII).
WeEnv §3's motivation analysis is very clean, with quantification in each of the three parts.
(1) Packaging is a dilemma (§3.2)
| Scheme | Approach | Cost |
|---|---|---|
| Dynamic assembly | Artifact excludes evolving components, installs at initialization | ~20 seconds extra per environment to install harness (~13% of initialization) |
| Static bundling | Full packaging into monolithic artifact, zero install at initialization | Combinatorial explosion: T tasks × H harness variants = T×H pre-packaged images |
The pain of static bundling is made very clear in the original: "Worse, the harness is not the only fast-evolving component; as described in §2, the evaluator is also updated routinely against newly discovered reward hacking. Each such change touches every affected image, which can take days and lengthens the experimentation cycle of agentic RL."
Note an implicit echo between two institutions here: WeChat says "the evaluator must be frequently updated to combat reward hacking"; DeepSeek uses an entire section (§6.4) to record the reward hacking behaviors they actually caught (forged chronus RPC, log searching, XFS_IOC_SWAPEXT bypassing access control). Two companies independently confirm the same thing: combating reward hacking is a continuous engineering effort requiring frequent environment iteration, not a one-time data cleaning.
(2) Initialization's I/O amplification (§3.3)
WeEnv traced all I/O to the rootfs during each task's execution (on OverlayFS, because partial writes first read the whole file, so measuring only reads covers all required bytes), obtaining a CDF:
"the median task reads merely 0.81% of its artifact, and even the 99th percentile reaches only 6.02%; no task in our trace touches more than 7%." "fetching the full image transfers over 100× the bytes that execution ever accesses."
(3) Provisioning's fixed quota (§3.4)
Profiling peak demand from 8k tasks of SWE-Smith-Python (Figure 4):
| Metric | p50 | p90 | p99 | Span |
|---|---|---|---|---|
| Peak CPU | 0.99 cores | 17.7 / 18 cores | 68.3 / 68 cores | ~70× |
| Peak memory | 0.15 GiB | 0.17 GiB | 7.0 GiB | — |
Memory is even more skewed: 92% of tasks peak below 0.2 GiB, barely above the 0.15 GiB median, then p99 jumps directly to 7.0 GiB. The original's arithmetic is striking:
"One sized for the median starves the tail, while one sized for the tail wastes resources on the majority, e.g., a 7 GiB quota would over-provision nine out of ten tasks by more than 35×."
And quota directly translates into latency (Figure 5): for the same resource-intensive task, under the minimum quota (2 cores / 8 GiB) rollout takes 310.5 s, while under the maximum quota (128 cores / 512 GiB) it takes only 201.6 s—tight quotas stretch rollout by 1.5×.
Why prediction is not possible (§4): WeEnv sampled 10 resource-intensive tasks aligned at their start to observe demand evolution (Figure 6). The conclusion is prediction is infeasible—demand has no clear pattern, the same task varies greatly across rollouts, and bursts are sudden: one task's CPU demand surged ~47× within one second. So WeEnv simply gives up prediction and switches to elastic adjustment.
Problem positioning is very precise (§4):
"The environment tooling of both virtual machines and containers already embraces a similar notion, the layer. ... However, layers serve only as packaging-stage units of caching and deduplication. At initialization, they must have been bound into a single monolithic artifact, one template or one image, before an environment can launch." "The fundamental gap is that layers cannot be located on their own. They exist only as entries in the artifact's manifest... Layers thus cannot be indexed apart from the enclosing artifact."
Key technical detail: Docker's registry protocol only accepts image references, not bare digests. So under the standard protocol, layer-only addressing is unavailable.
WeEnv's answer (§6.1):
harness:2.0, task1:1.0, base-env:3.0), independently versioned and published;The original points out the only essential difference from a conventional artifact:
"the artifact freezes its layer stack at packaging time, whereas the plan assembles it at initialization from independently published groups, so the same groups recompose into different environments without repackaging."
And the cost of composition is extremely low: "it manipulates only metadata, resolving cached references, concatenating digest lists, and issuing one mount, while no layer content is copied, unpacked, or executed."
Packaging granularity cut by change rate (§6.2). This is the design guideline itself: "The guideline is to draw group boundaries along change rates." So it cuts into four groups:
| Group | Content | Change rate |
|---|---|---|
| harness group | Harness coordinating model and environment (Claude Code / Codex, etc.) | Highest, continuously upgraded |
| evaluator group | Scoring and reward computation | High, continuously patched against reward hacking |
| task group | Initial state per task (git checkout, faulty patch) | Medium |
| base group | OS, language runtimes, common tools | Low, shared by all tasks |
So a harness upgrade or evaluator patch publishes only a small group, touches none of the T task groups, and the change takes effect in every environment composed afterward.
One more layer: repository-granularity packaging (§6.2 end). The original observes a very strong structure:
"many RL tasks share one code repository and differ only in a small mutation, such as checking out a particular commit or applying a faulty patch. For example, the 39,471 tasks of SWE-Smith-Python span merely 131 repositories."
So task is packaged at repository granularity (all tasks in the same repository share), each task only needs a cheap mutation applied at initialization—measured median 0.43 seconds (1024 initialization measurements), compared to ~20 seconds to install harness.
Results (§9.3, Table 1):
| Dataset | Dynamic assembly | Static bundling | WeEnv |
|---|---|---|---|
| SWE-Smith-Python | 39,471 | 39,471·H | 133 + H |
| SWE-Smith-C++ | 5,123 | 5,123·H | 71 + H |
| SWE-Smith-Go | 1,629 | 1,629·H | 21 + H |
| SWE-Smith-Java | 6,704 | 6,704·H | 61 + H |
| SWE-Smith-JS | 6,073 | 6,073·H | 36 + H |
| SWE-Smith-Rust | 5,311 | 5,311·H | 41 + H |
| SWE-Smith-TS | 5,032 | 5,032·H | 32 + H |
The general formula is B + 2 + H (B repository-level base images + 1 base tool + 1 evaluator + 1 per harness). From 39,471 (or ×H) down to 133 + H, about two orders of magnitude. More critically, the change cost: under static bundling, changing harness or evaluator requires touching "up to T images of the entire dataset"; under WeEnv, exactly one layer group is republished.
Data flow (§7.1): To intercept file I/O, WeEnv mounts the rootfs into the environment via FUSE. When a task accesses a file, the kernel forwards I/O to envlet's FUSE handler, which uses layer metadata to translate it into a request for "a specified layer + specified byte range." Requests are served in three tiers:
envlet cache (shared by all environments it hosts)
└─ miss → node-level cache (one per node, shared with peer nodes)
└─ miss → remote registry
Only reads enter this path—all writes fall to the local writable layer; partial writes to files backed by read-only layers trigger copy-up, which first reads the whole file, so it also appears as a read.
There is also a best-effort background fill: once an environment starts, it traverses the artifact chunk by chunk, skipping chunks already brought in by on-demand fetching, keeping completion off the critical path.
Why the layer format must be rebuilt (§7.2). Two requirements, Docker's conventional layer format satisfies neither:
| Requirement | Conventional format (gzip-compressed tar stream) | Consequence |
|---|---|---|
| Metadata independently pullable | No central directory; each file header is adjacent to its data, interleaved with content | Recovering metadata requires walking the entire layer |
| Arbitrary byte ranges independently readable | gzip compresses an archive into a single sequential stream; later bytes reference earlier ones | Reading any byte means decompressing all content before it |
WeEnv's layer format:
Measured gains (§9.4):
| Metric | On-demand off (full download then start) | WeEnv on-demand | Improvement |
|---|---|---|---|
| Median initialization | 32.5 s | 10.6 s | 3.1× |
| P95 | 83.5 s | 35.1 s | 2.4× |
| Worst case | 208 s (layer must come from remote registry) | 47 s | 4.4× |
| 1st RL iteration (all cold) | 1,845 s | 763 s | 2.4× |
| Sum of all iterations | 8,092 s | 5,988 s | 1.35× |
The original's qualitative conclusion is key: "on-demand fetching removes the artifact transfer from the critical path regardless of the cache state"—initialization latency no longer depends on artifact size.
Monitoring (§8.1): Every environment starts from the same small quota: 2 cores / 8 GiB. envlet manages resources via cgroup, and monitors through the same interface—reading cpu.stat / memory.stat every 200 ms. It emphasizes that monitoring itself is extremely cheap: nothing runs inside the environment, each environment round takes tens of microseconds, and at 200 ms intervals uses less than 0.1% of a single core.
Principle One: Scale up aggressively, scale down conservatively (§8.2). The reason is that demand can rise 47× in one second, and an incorrect downscale directly kills the environment and loses accumulated multi-turn interactions.
| Pressure level | Criterion | Action |
|---|---|---|
| Normal pressure | CPU: quota utilization ≥85%, or 10% of scheduling periods throttled; memory: working set ≥60% of quota, or cgroup reports reclaim events/pressure | Double quota |
| Severe pressure | CPU: throttling exceeds 80%; memory: OOM emergency | Quota immediately ×4 |
| Downscale (CPU) | Utilization persistently below 30% for 10 seconds and throttling negligible | Halve quota |
| Downscale (memory) | Working set persistently below 20% of quota for 60 seconds | Halve quota |
Principle Two: Scale-up must not starve admission (§8.2). This scenario is designed very realistically: because some tasks must be evaluated in independent clean environments, a batch is admitted in two waves—first execution environments, then evaluation environments after execution completes. The two waves overlap on the same node; if scale-up of running environments consumes all admission capacity, evaluation environments cannot start. Two constraints:
Principle Three: Failures must be explicit, not silently degraded (§8.2).
"A trajectory silently corrupted by starvation is worse than a lost one, as it feeds wrong signals into training. When an OOM emergency persists with no tier left to raise, the envlet terminates the task and reports the cause, so that the RL framework re-samples the task instead of learning from a starved execution."
This is the most copy-worthy design judgment in the entire article. It elevates "resource allocation failure" from a performance issue to a training data correctness issue: rather lose a trajectory than let a starved trajectory enter the training set. The same thinking appears again in WeEnv's architecture—evaluation must run in an independent clean environment, "so that rollout-side modifications cannot contaminate the reward computation, guarding against reward hacking."
Measured gains (§9.5):
| Metric | Fixed quota | Elastic provisioning | Improvement |
|---|---|---|---|
| Single-task rollout | Slowest 2 hit 1,800 s agent timeout | Same task ~40 s | 45× |
| Executions saving >100 s | — | 61 | — |
| 3rd / 5th / 8th iteration | 1,949 / 1,773 / 1,981 s | 542 / 806 / 534 s | up to 3.7× |
| Sum of all rollout latencies | 8,968 s | 5,988 s | 1.50× |
The original also proactively discloses counterexamples: a few executions on the right end are actually slower under elastic mode, because "their extra time is spent in the solving phase, where the agent takes a different, longer trajectory, a stochastic effect of agent sampling rather than of the resource allocation." This clarification is very well done.
Quota timeline (§9.5, Figure 12) gives a complete example: CPU starts at 2 cores, one burst during solving jumps from 4 directly to 16, then gradually halves back to 2 as the burst subsides (released cores "serve other environments on the same node for most of the solving phase"); during evaluation another burst occurs, quota climbs past 32 up to the single-environment cap 64. Memory is a different story: it stays at the initial 8 GiB throughout the solving phase (that phase is dominated by model calls, memory consumption is small), until the evaluator needs it and doubles twice to 32 GiB. When the task completes the environment is torn down, and quota immediately drops to zero, without waiting for conservative downscale.
The value of this example is that it proves the necessity of elastic quotas: "the same environment holds widely different quotas at different moments, which no fixed allocation could match."
(§9.2) Same workload run on four backends:
| Metric | E2B | Docker | AgentENV | WeEnv |
|---|---|---|---|---|
| Env init share of iteration time | 53.4% | 39.5% | — | 9.1% |
| Task exec share of iteration time | 20.2% | 44.0% | — | 74.0% |
| Median initialization | 150.6 s | 111.0 s | 59.7 s | 10.6 s |
The three baselines' initialization shares range from 28.9%–53.4%; WeEnv compresses it to 9.1%, with task execution rising to 74.0%—the original summarizes: "the iteration finally spends its time on solving the tasks rather than on preparing the environments."
Median initialization 10.6 s vs 59.7 / 111.0 / 150.6 s, i.e., 5.6–14.2×. The distribution shape is also worth noting: Docker's curve is bimodal—38% of initializations hit a warm container on the node and complete within 30 s, the rest pay 100–300 s for full cold creation; while WeEnv "keeps the entire distribution low."
Robustness experiments (this is where the paper's quality shows; sensitivity was tested in three dimensions):
| Variation | WeEnv | E2B | Docker | AgentENV |
|---|---|---|---|---|
| Switch harness to Codex, P50 task time | 108 s | 260 s | 228 s | 178 s |
| Same, median initialization | 13.7 s | 160.2 s | 126.8 s | 87.3 s |
| Qwen3-14B + 8 servers, P50 task time | 188 s | 288 s | 253 s | 207 s |
| Same, median initialization | 13.9 s | 150.7 s | 108.2 s | 60.4 s |
When switching datasets (Rust / C++ / Go / JS), WeEnv is also fastest: 66.9 / 131.5 / 78.9 / 105.3 s. The original's explanation of the differences is apt: WeEnv's savings come mainly from the environment side, roughly constant per task, so the relatively faster the task itself executes, the greater the relative gain (up to 3.4× on Rust, Go), while C++ is compilation-intensive and task time itself dominates, narrowing the advantage over the strongest baseline to 1.2×—but still ahead.
WeEnv explicitly compares AgentENV's storage design in §9.2; both criticisms point to architecture rather than implementation:
(1) Block-level conversion is not cache-friendly.
"it converts file-level layers into block-level ext4 layers with overlaybd, and the conversion is not cache-friendly: the same layer may map to different blocks across images, so the converted blocks cannot be shared by digest, an issue also observed in file-oriented image services [31]."
Compare its own approach: "WeEnv instead caches at the layer level, where a cached layer is reused wherever its digest matches, regardless of the enclosing image."
This criticism is structural: file-level layers have content-addressed digests, naturally globally deduplicable and shareable; once converted to block-level ext4, "the same upper-layer content" may land on different blocks in different images, breaking cross-image digest-based reuse. This is an inherent cost of the "block-level route," not an AgentENV implementation bug.
(2) Narrow optimization scope.
"AgentENV optimizes solely the initialization, whereas WeEnv spans the full lifecycle, rethinking packaging and resource provisioning as well."
Against the "three narrow gates" diagram in Section I: this is exactly the inevitable result of AgentENV betting on "startup + state"—it puts almost all engineering effort into initialization and snapshot/fork; packaging and provisioning are outside its scope.
T×H comparison is slightly simplified: static bundling's T×H is an upper bound (in reality deduplication may not yield that many), but "changing evaluator once touches all affected images" is independent of the upper bound, so it does not affect the argument.AgentENV comes from Moonshot AI, code under GitHub org kvcache-ai. The repository states:
"AgentENV (AENV) is a platform for running agent environments at scale, powering agentic RL training for Kimi K3." (README:22)
Verifiable repository facts (local checkout):
| Item | Value | Location |
|---|---|---|
| License | MIT | LICENSE:1 (Copyright (c) 2026 AgentENV) |
| First open-source commit | 2026-07-25 feat: Initial open-source release of AgentENV | git log --reverse |
| Total commits | 235 (as of local checkout) | git rev-list --count HEAD |
| Production model claim | Kimi K3, external link to github.com/MoonshotAI/Kimi-K3 technical report | README:22, 28 |
| Storage crate metadata | name = "overlaybd", version = "0.1.0", no repository / authors / license fields | storage/overlaybd/Cargo.toml:1-4 |
| Disclosure of DeepSeek | Searching all *.md for deepseek / dsec zero hits | See §1.2 |
Its positioning is fundamentally different from the other two: DSec and WeEnv are "systems in papers"; AgentENV is "something you can directly install" (install.sh / aenv CLI / E2B SDK compatible). So when evaluating it, "what is in the code" is more credible than "what the README says"—below I will strictly separate the two.
AgentENV's own architecture document states its focus bluntly (docs/src/internals/architecture.md:3):
"AgentENV runs AI agents inside isolated, snapshot-capable Firecracker microVMs. Its core is a storage subsystem that provides layered block devices mountable into VMs and ublk-backed memory snapshot restore. The system also includes a per-node orchestrator managing sandbox lifecycle, and a distributed control plane (gateway + scheduler) for multi-node routing."
Three-layer structure:
AgentENV Node
├── API (Axum) ── E2B compatible + reverse proxy
├── Orchestrator ── lifecycle state machine (Creating/Running/Pausing/Paused/Resuming/Killing)
├── Firecracker VM
│ ├── /dev/vda rootfs ── ublk /dev/ublkbN ── overlaybd (upper writable + layer 0..N read-only)
│ ├── /dev/vdb extra drive ── ublk ── overlaybd
│ └── VM memory ── ublk read-only memory block device ── overlaybd mem layers
│ (multiple sandboxes of the same snapshot share this one device via reference counting)
└── Cluster side: Gateway(:8080) ── gRPC ── Scheduler(:9090)
Note the shape of the memory path: the memory snapshot is not a file, but another ublk block device, using the same overlaybd layering logic as rootfs. This is where AgentENV differs most from the other two—it turns "memory" into a block device that can be layered, read on demand, and shared by multiple sandboxes.
overlaybd's format (architecture.md:47-67): not a simple layered tar, but an LSMT (Log Structured Merge Tree) format:
HeaderTrailer (magic LSMT\0\1\2, UUID, flags, index/data offsets) and a set of DiskSegmentMapping entries, each 16 bytes bit-packed: 50-bit offset, 14-bit length, 55-bit physical offset, zeroed flag, layer tag.ImageFile searches each layer's segment index top-down, the first layer containing the block range mapping supplies the data; ranges not mapped by upper naturally fall to lower.VirtualFile trait): LocalFile (io_uring pread/pwrite, optional O_DIRECT), registryfs_v2 (OCI registry remote layer download), tar, plus optional decompression block cache.ublk device lifecycle (architecture.md:69-85):
UVMUblkCtrlBuilder sends ADD to /dev/ublk-control via io_uring's UringCmd;/dev/ublkcN (control) and /dev/ublkbN (block);AsyncIoRing and slab-allocated I/O slots;ublksrv_io_desc array, processed asynchronously in user space;AutoRegBuffer (sparse buffer table for zero-copy, kernel 6.8+) or traditional UserBuffer.A detail only known from hitting the pit (architecture.md:92): ublk devices are uniformly held by a resident daemon uvm-ublk-daemon, with node side sending RPC over Unix domain socket. The reason is very practical:
"All remote image I/O (OSS/registry reads) is dispatched from the per-queue runtimes onto a dedicated per-
ImageServiceremote-io runtime viaRuntimeDispatchFile, because opendal/reqwest pin pooled HTTP connection tasks to whatever runtime polls them and per-device runtimes are dropped on device teardown."
That is: tokio's HTTP connection pool pins tasks to "the runtime that first polled them," and per-device runtimes are dropped on device teardown—if remote I/O ran directly on per-queue runtimes, when the device is destroyed, tasks in the connection pool become orphans. So all remote I/O is dispatched to an independent, resident remote-io runtime (controlled by ublk.overlaybd.remote_io_workers, default 4, see config/default.toml:313). This is a pure runtime ownership problem, not an algorithm problem, and it is exactly the dividing line between "it runs" and "it explodes after running for a while."
This is the deepest part of AgentENV among the three, and where it puts all its engineering effort.
(a) pause is just a "seal layer + reopen upper" action (architecture.md:65)
"
ImageFile::create_snapshot_and_restack()is the primary pause path. It seals the live upper layer viaLSMTFile::close_seal_and_reopen()so the upper becomes the newest lower layer, then reopens a fresh writable upper in place."
Note the consistency of this design: after sealing, the original upper becomes the youngest read-only layer, and a new writable upper opens in place—so after pause the environment can immediately continue running, while the entire filesystem increment it just accumulated has become a reusable lower layer. Snapshot is not an "export" action, but a rearrangement of layer stack pointers.
(b) What a snapshot retains (docs/src/concepts/snapshots/index.md:13-24)
(c) fork is a first-class API, not a wrapper around "copy snapshot"
The route actually exists in the OpenAPI contract and generated code: POST /sandboxes/{sandboxID}/fork (src/api/openapi.yml:2045, src/api/generated/src/server/mod.rs:73). The contract semantics are clear:
"Number of forked sandboxes to create. All forks boot from the same snapshot, so the snapshot is captured once regardless of count. Each fork succeeds or fails independently; the outcome of each is reported in its entry of the response list." (
src/api/openapi.yml:836-838)
Several judgments worth recording at the implementation layer (src/orchestrator/service.rs:629-860, src/sandbox/firecracker/sandbox.rs:302-500):
FirecrackerSandbox::pause(self) obtains the snapshot then immediately FirecrackerSandbox::resume(self), then uses futures::future::join_all to concurrently bring up all children (src/sandbox/firecracker/sandbox.rs:452-459, 488-502). The source sandbox is paused only for an instant; there is no "copy a memory image" action.snapshot_config_for_fork() generates replacement drive configs for each child (sandbox.rs:302), and validates "fork replacement drives must refer to launch-time volume drives" (sandbox.rs:474).fork_sandbox_keeps_successful_siblings_when_one_start_fails (src/orchestrator/tests.rs:5345) locks this behavior—one child failing to start must not drag down already-started siblings.CLAUDE.md:195:| Mode | Behavior at fork | Location |
|---|---|---|
exclusive (default) | Each child creates an independent COW subvolume, and drive replacement relationships are registered | src/api/impls/sandbox.rs:221-243; enum default see src/volume.rs:36-41 |
ro | Reuses the same volume id (read-only, multi-owner safe) | src/api/impls/sandbox.rs:245-248 |
The underlying reservation semantics are consistent: exclusive is single-owner reservation, conflicts error directly (src/volume.rs:591-600); ro records multiple owners (:579-589). Documentation matches code (docs/src/concepts/volumes/sandbox-forks-and-snapshots.md:12-16).
prepare_volume_fork_specs() creates volume-fork-{uuid} subvolumes; any step failing triggers cleanup_fork_volume_children() to clean up already-created ones (src/api/impls/sandbox.rs:191-255).Running; terminal errors require tearing down the source sandbox (tests fork_sandbox_recoverable_failure_cleans_up_metrics vs fork_sandbox_terminal_failure_removes_source_and_cleans_up_metrics). This matches the constraint in CLAUDE.md: "Recoverable capture/fork errors restore Running; terminal errors require teardown."(d) Memory snapshot bypasses userfaultfd, goes through block device—this is the highest technical content here
architecture.md:112-120 records a clear architectural turn:
"Memory snapshot restore uses ublk-backed overlaybd devices rather than userfaultfd. On resume, a read-only ublk device is created from the stacked memory overlaybd layers and passed to Firecracker as a
BackendType::Filememory backend. Firecracker mmaps the block device and COWs pages into anonymous memory on first write, so the underlying device is never modified."
Three consequences, each with practical value:
→ This directly serves the scenario of "instantly bringing up a large batch of environments from one template," and sharing happens at the host page cache layer, requiring no deduplication algorithm.
process_vm_readv to read selected memory, directly creates an OverlayBD memory layer; and stacks the parent layers of previous snapshots to form a complete layered memory image. → Sandboxes that repeatedly pause/resume pay only the increment each time.But increments are not free. Layers accumulate continuously, so the code configures a layer budget: DEFAULT_MAX_OVERLAYBD_SNAPSHOT_LAYERS: usize = 32 (src/sandbox/firecracker/overlaybd_snapshot.rs:41, comment says it is the "base budget for runtime-owned snapshot lowers before merging"). Exceeding triggers compaction to collapse multiple layers.
This is the half most easily missed in "incremental snapshot" designs: DSec solves the same problem on the container side with offline layer collapsing (~3 GB threshold) (§3.5); AgentENV uses a 32-layer budget at runtime to contain it. Both independently discovered: a layering scheme must be paired with an answer to "what if there are too many layers," otherwise read amplification and mount overhead gradually eat the gains.
Corroborating the thoroughness of this turn: storage/uffd-core/ remains in the repository, but is excluded from the workspace build (architecture.md:120: "retained for reference but excluded from the workspace build"). That is, the userfaultfd path is abandoned, not optional.
(e) warm pool: moving two things off the resume critical path in advance
architecture.md:142:
"Snapshot resume can also use
[pool.firecracker]to pre-spawn(network slot, Firecracker process)pairs. A warm entry transfers its network slot, process, and Firecracker CWD to the resumed sandbox, which avoids the spawn and API-socket wait in the resume critical path."
Config defaults are on (config/default.toml:338-344): pool.firecracker.enabled = true, startup_prewarm = true, fill_concurrency = 4. There is also [pool.block] reusing the OverlayBD ublk device pool (default.toml:331-336), and shared watermarks low_watermark = 2 / high_watermark = 64 (default.toml:321-325).
This is a very typical piece of "startup latency engineering": latency is not saved in the algorithm, but work is moved to idle periods (pre-spawn processes, pre-warm network slots, pre-create block devices) and then ownership is transferred on the critical path.
AgentENV's guest kernel command line (config/default.toml:33) contains this string:
page_reporting.page_reporting_order=6
damon_reclaim.enabled=Y
damon_reclaim.min_age=100000
damon_reclaim.quota_ms=20
damon_reclaim.quota_sz=1073741824
damon_reclaim.quota_reset_interval_ms=500
damon_reclaim.wmarks_high=990
damon_reclaim.wmarks_mid=990
damon_reclaim.wmarks_low=200
damon_reclaim.skip_anon=Y
damon_reclaim.wmarks_interval=1000000
And balloon-side config—src/sandbox/firecracker/instance.rs:417-434 explicitly enables free page reporting:
pub async fn set_balloon(&self) -> Result<()> {
let balloon = Balloon {
...
free_page_hinting: None,
free_page_reporting: Some(true),
};
The config file comments also specifically explain the purpose of this combination (config/default.toml:31-33: "Editing it replaces the shipped boot and DAMON reclaim settings"), while src/sandbox/firecracker/config.rs:38-39 explains the meaning of each wmarks item.
Here appears a technical echo that must be pointed out: DSec §5.2 uses the same set of kernel mechanisms—"We therefore combine DAMON (Data Access MONitor) with virtio-balloon free-page reporting", and empirically proves DAMON + balloon FPR reduces time-integrated memory consumption by 21.2%.
The two take different values on the same parameter, which is interesting:
| DSec §5.2 | AgentENV default.toml:33 | |
|---|---|---|
| Page reporting granularity | "operates on order-9 pages, corresponding to 2 MiB regions with 4 KiB base pages" (default, adjustable) | page_reporting.page_reporting_order=6 (256 KiB) |
| DAMON cold page reclaim | Enabled (with balloon FPR) | damon_reclaim.enabled=Y, skip_anon=Y |
Convert: order-9 = 2⁹ × 4 KiB = 2 MiB, order-6 = 2⁶ × 4 KiB = 256 KiB—AgentENV makes reporting granularity 8× finer. For microVM scenarios this is reasonable: finer granularity means more free pages can be assembled for reporting and reclamation is more precise, but scan overhead is also greater. This is a typical engineering choice "papers won't write, only known after tuning."
It must be clear: identical kernel parameters alone cannot prove code homology (DAMON + FPR is a combination any team doing high-density Firecracker deployment would encounter). But together with the storage component relationship already disclosed in the DSec paper in §1.2, it forms a coherent picture: these two use the same technical vocabulary and the same batch of kernel features at the two "kernel boundary layers" of block storage and memory reclaim. Their platform layers (scheduling, lifecycle, security) remain completely different.
External interface directly compatible with E2B:
POST /sandboxes create
GET /sandboxes list
GET /sandboxes/{id} metadata
DELETE /sandboxes/{id} delete
POST /sandboxes/{id}/pause pause (snapshot)
POST /sandboxes/{id}/resume resume from snapshot
POST /sandboxes/{id}/fork fork (N children)
GET /nodes, /nodes/{id} node view
ANY /proxy, /proxy/{path} HTTP/SSE/WebSocket reverse proxy into sandbox
(docs/src/internals/architecture.md:181-192; fork is not in that list, it is defined by the contract: src/api/openapi.yml:2045)
How "E2B compatible" is specifically achieved: src/api/openapi.yml never mentions "e2b"—compatibility is not through an E2B-specific spec, but through aligning route shapes, request headers, and ID formats:
| Compatibility point | Location |
|---|---|
Route header aliases e2b-sandbox-id / e2b-sandbox-port | src/api/proxy.rs:87-92 (parsing at :818, :825, stripping before forwarding at :973, :975) |
Traffic token header e2b-traffic-access-token | src/api/impls/auth.rs:14 |
API key prefix e2b_ | src/api_key.rs:20 |
| MMDS field names / template step error shapes / default ready command aligned with E2B | src/sandbox/firecracker/mmds.rs:68, src/template/step_executor.rs:473, src/template/runner.rs:28 |
And it clearly states what is not compatible (docs/src/integration/e2b.md:96-97): "AgentENV supports the TypeScript SDK's volume create, list, and delete operations... The E2B SDK's direct volume content API is not supported."—A compatibility layer proactively marking its gaps is far more useful than claiming "fully compatible."
There is also [custom_extension] lifecycle hooks (default.toml:131-136), [sandbox_proxy].domains hostname-based routing, and the aenv CLI (Auth / Pull / Build / Start / Exec / Connect (alias cn) / Pause / Resume / List (ls) / Delete (rm) / Timeout / Snapshot (snap) / Template / Volume, see crates/aenv/src/main.rs:20-65).
The control plane is an independent Go module (services/):
Client ──HTTP──> Gateway(:8080) ──gRPC──> Scheduler(:9090)
│
└──proxy HTTP──> Node A / Node B (:8000)
architecture.md:209 lists 8 (Schedule, LookupNode, RecordAssignment, Heartbeat, ListObservedNodes, ListP2pPeers, GetNode, UnregisterNode), but the contract actually has 13 (services/api/proto/scheduler.proto:7-21, plus ListNodes, ReportSandboxEvent, and three P2P artifact RPCs)—the documentation's RPC list is incomplete; strategies round_robin (default) / random; discovery modes static or kubernetes (EndpointSlice watch).RecordAssignment immediately creates the binding; runtime heartbeats carry the node's full sandbox ID roster, and the scheduler deletes bindings not in the roster based on the roster; binding_ttl is route freshness, not a copy of sandbox timeout.Regarding "where state lives," there is a common misreading that must be corrected for AgentENV. architecture.md:223 says:
"Limitations: All bindings are in-memory (lost on scheduler restart). After a scheduler restart, bindings are rebuilt from new sandbox creations plus the next heartbeat roster from each runtime node... but binding persistence is still not replicated."
This sentence is too broad; copying it verbatim underestimates this control plane. The actual situation is:
| Fact | Location | |
|---|---|---|
| Default | Bindings in memory (NewInMemoryBindingStore) | services/scheduler/cmd/main.go:165 |
| But optional Redis backend | NewRedisBindingStore, Lua scripts for atomic writes | services/scheduler/internal/redis_store.go:28, 103-113 |
| And documented HA shape | One primary (with scheduler.redis_addr) + multiple --query-only replicas, the latter serving only LookupNode from Redis | services/README.md:293 |
| Still a real gap | "Artifact store state is still in-memory and is not covered by this HA mode." | services/README.md:293 |
So the accurate statement is: AgentENV's control plane has an optional Redis binding store and query-only replica scheme, but artifact indexing remains only in memory under HA mode; the sentence in architecture.md:223 is too broad or even outdated.
Even so, this control plane's positioning is clear: AgentENV's focus is on in-node storage and snapshots; the control plane only handles "routing requests to the correct node." This looks similar to DSec's repeatedly emphasized design orientation that "placement engine and watcher need no persistent state," but the paths are opposite: DSec is actively designed to be stateless and validates with periodic IaC rebuilds; AgentENV is stateful by default first, then adds an optional Redis backend, and explicitly admits HA does not cover artifact indexing.
AgentENV has no accompanying paper, so its performance claims can only come from the README. I distinguished "code-supported" from "text-only" item by item:
| Claim | Source | Code-level evidence | Verdict |
|---|---|---|---|
| boot/resume < 50 ms, pause < 100 ms | README:29 (same sentence in docs/src/getting-started/overview.md:9) | Has benchmark suite (crates/benchmarks/benches/snapshot_benchmark.rs: snapshot_resume, snapshot_resume_cold, concurrent_resume, concurrency CONCURRENCY = 50), but only times, no thresholds, and no result files in repo; mechanism side has warm pool (src/sandbox/firecracker/pool.rs comment: "entry to skip process spawn and API socket polling on the critical path") + pool.firecracker.startup_prewarm = true | ⚠️ Claimed value, no threshold/no results |
| 9.6× memory overcommit ratio (production) | README:31 | Mechanisms complete (balloon FPR + DAMON + memory device reference-count sharing + control plane allows >100% allocation: services/shared/config/config.go:46-48 max_memory_allocated_percent "can exceed 100 (overcommit)"); but the string 9.6 appears elsewhere in this repo only in unrelated crate version numbers in Cargo.lock, with no measurement or config item corresponding to it | ⚠️ Specific value unsupported; and docs/src/getting-started/overview.md:11 deletes "9.6x" in the same paragraph |
| snapshot < 100 ms, even with heavy disk modification | README:30 | Incremental capture mechanism real (src/sandbox/firecracker/sandbox.rs:1345-1351: state-only diff snapshot + get_dirty_memory_ranges()); heavy-disk scenario has bench variant (bench_snapshot_creation_1gdisk) | ⚠️ No 100 ms assertion |
| 1.5 million images in production | README:28 | External link MoonshotAI/Kimi-K3 technical report, not in this repo | ⚠️ Cannot verify locally |
| native snapshot / fork | README:30 | ✅ Complete implementation: pause restack, POST /fork, per-child independent results, volume COW subvolumes, graded failure semantics, dedicated tests | ✅ Verifiable |
| on-demand loading of OCI images | README:28 | ✅ image.resolver.convert_standard_oci = true (default.toml:76), registryfs_v2 backend, image.cache.remote_blocks.max_size_gb = 100 (default.toml:95-97) | ✅ Verifiable |
| local disk as bounded cache, hot retained, cold evicted | README:28 | ✅ image.cache.capacity_gb = 100 + GC (interval 1800 s, min free 600 s, high watermark 0.95 / low watermark 0.70, default.toml:92-111) | ✅ Verifiable |
| P2P acceleration | default.toml:113-119 | ⚠️ Disabled by default (enabled = false), and comment self-describes "EXPERIMENTAL: P2P has not been tested in production. Keep it disabled in production unless the deployment accepts that operational risk."; another factory config self-contradiction: [snapshot].p2p_enabled = true (:228) while [p2p].enabled = false (:117), the former does not take effect due to forced downgrade in src/p2p/config.rs:36-40 | ⚠️ Self-declared unverified |
The two most noteworthy points in this table:
9.6x only appears in the README, and is deleted in the corresponding docs paragraph. (README:31 vs docs/src/getting-started/overview.md:11: the latter retains "Memory ballooning returns reclaimable guest memory to the host," but rewrites "achieving a 9.6x memory overcommit ratio in production" to "sustaining high overcommit.") This inconsistency is an observable difference within the documentation—I cannot determine whether it is deliberate or an oversight, but citing this number should note that the source is README only.[template_build] only builds template images with BuildKit (builder_image = "moby/buildkit:v0.33.0", max_concurrent_builds = 4, cache_size_mb = 65536, default.toml:376-385), and has a per-node concurrency limit (returns 429 when exceeded). WeEnv's criticism holds here: AgentENV is not responsible for "how environments are packaged and maintained"; it assumes the artifact already exists, and focuses on turning the artifact into a fast sandbox.docs/src/concepts/volumes/index.md:3-6: "The volume feature is still in beta. For workloads that can use a drive supplied only when cold-starting a sandbox, the attachedDrives feature on POST /sandboxes-cold is more extensively tested.")—and fork's volume semantics happen to depend on volume; the userfaultfd path is abandoned and removed from the workspace.docs/src/internals/sandbox-testing.md records defaults that are all stale (:313 says socket_poll_ms default 20 ms, code is 1; :319 says machine default "128 MiB RAM and 1 vCPU," code is 1024 MiB / 2 vCPU; :321-322 says envd 30 s / 10 ms, code is 60 / 3), while docs/src/configuration/reference.md matches code; architecture.md:63 says compression is "zstd (level 3)," while the factory release compression default is lz4 (default.toml:246, both supported); architecture.md:223's "all bindings in memory" is too broad (actually there is an optional Redis backend, see §5.6). These do not affect functionality, but mean: citing any specific value from AgentENV docs must be checked against code.sandbox.rs:474), and volume fork goes through COW subvolumes—fork is not "unlimited free copying of arbitrary state," but a first-class operation with clear boundaries.| Dimension | DSec (DeepSeek + Tsinghua) | WeEnv (Tencent WeChat) | AgentENV (Moonshot) |
|---|---|---|---|
| Material form | Technical report (31 pages) | Paper (16 pages) | Open-source codebase |
| Isolation base | FnCall / container / Firecracker microVM / QEMU full VM (4 kinds) | Container (only) | Firecracker microVM (only) |
| Container wrapped in VM? | Yes (containers run inside QEMU/libvirt VMs) | No | N/A |
| Environment packaging unit | base + workspace + toolkit (EROFS layers) | layer group + environment plan | Does not do packaging |
| Who builds environments | Agents themselves online (pack_diff) | Offline cut by change rate + repository-granularity packaging | External (BuildKit builds templates) |
| On-demand loading | Containers: EROFS multi-device; microVMs: overlaybd + ublk, 256 KiB chunk + local second-level cache | FUSE + custom layer format (tail TOC + contiguous payload) | overlaybd (LSMT) + ublk, remote blocks cache |
| Content addressing granularity | File-level (EROFS) / block-level (overlaybd) | File-level (layer digest reusable across all images) | Block-level (LSMT segment) |
| Density measures | virtio-pmem+DAX, DAMON+FPR, CPU QoS (core scheduling) | Elastic provisioning (200 ms cgroup sampling) | balloon FPR + DAMON + memory block device reference-count sharing |
| Snapshot / fork | pause/resume (container docker pause; microVM snapshot+kill process) | Not a core mechanism | First-class: restack pause, POST /fork, memory device sharing |
| Coordination with trainer | Deep co-design (rollout moved out of preemptible GPU pool) | Robustness only (clean evaluation environment prevents contamination) | None (positioned as general platform) |
| Adversary model | Most complete (AppArmor + eBPF network allowlist + case postmortems) | Architectural (independent clean evaluation environment) | Not public |
| External interface | libdsec (Python SDK) | Integrated with Slime | E2B-compatible HTTP API + aenv CLI |
| Scale evidence | 160 nodes/scale unit, 3M sandboxes/day, 380K concurrency, >5,000 creations/s | "Deployed at WeChat," scale undisclosed; evaluation 4–6 H20 | README claims 1.5M images; only third-party measurement is weakened |
| Performance measurement credibility | Self-measured, with ablations, 10-node experimental cluster | Self-measured, with 3-dimension sensitivity analysis, and actually measured three competitors | No independent measurement |
Reading the summary table vertically reveals something easily overlooked: these three have hardly ever clashed head-on on the same technical decision.
The orthogonality of the three bets shows in several concrete places:
| Question | DSec's answer | WeEnv's answer | AgentENV's answer |
|---|---|---|---|
| Where does the environment come from? | Agents build online | Offline cut by change rate | Doesn't care |
| How does the environment reach the node? | EROFS/overlaybd on-demand from 3FS | Custom format on-demand from registry | overlaybd on-demand from registry/OSS |
| After it's running? | Coordinate pause/resume with trainer | Elastically adjust CPU/memory quota | Snapshot, fork, share memory pages |
| Who is most dangerous? | The model itself (reward hacking) | Contaminated evaluation signals | Not stated |
So "who is better" is the wrong question. The right question is: which cut you need to make, and whose range covers it.
A coarse-grained selection guide:
| If your constraint is… | Closer answer |
|---|---|
| Security-sensitive tasks / full OS environments / Android / graphics / GPU kernel | DSec (four-backend spectrum, no alternative) |
| SWE-type container-reachable workload, pain is "iteration time eaten by environments" | WeEnv (full lifecycle + elastic quota) |
| Need environment state portability, fork branching, or an E2B-compatible self-hosted base | AgentENV |
| Need preemption coordination with RL trainer (GPU jobs frequently preempted) | DSec (the only one that did this) |
This is, I think, the most memorable point of the entire article. All three say "environment," but their denotations of the word "environment" are fundamentally different:
| What "environment" is | Implied responsibility boundary | |
|---|---|---|
| DSec | A programmable artifact (base + workspace + toolkit, agent can pack_diff anytime) | Environment is something produced at runtime; the platform must support online building |
| WeEnv | A software supply chain requiring governance (4 layer groups with different change rates) | Environment is a release artifact requiring continuous maintenance; the platform must manage its full lifecycle |
| AgentENV | An input (an existing OCI image) | Environment is a given premise; the platform only turns it into a fast sandbox |
This definitional difference directly determines whether WeEnv's criticism holds. WeEnv says AgentENV "optimizes solely the initialization"—from AgentENV's own boundary definition, this is not a defect, but its definition itself: packaging is not its responsibility. Conversely, DSec's pack_diff is possible precisely because it does not hand packaging to an external pipeline.
One-sentence summary of the three's division of labor:
DSec manages "how environments are created"; WeEnv manages "how environments are maintained and scheduled"; AgentENV manages "how environments are started and copied."
The three can be stacked, not substituted—which also explains why DSec would open-source its storage layer into AgentENV's repository: at the "block device path" layer, the two share; above it, each builds its own tower.
In the entire article, only one set of data puts products from three different institutions on the same experimental bench—WeEnv §9. It is the only public source for understanding AgentENV's relative performance. Therefore its qualifications must be itemized, otherwise it is easily misused.
| Metric (SWE-Smith-Python, Qwen3-8B, Claude Code) | E2B | Docker | AgentENV | WeEnv |
|---|---|---|---|---|
| Median environment initialization | 150.6 s | 111.0 s | 59.7 s | 10.6 s |
| Median initialization after switching to Codex harness | 160.2 s | 126.8 s | 87.3 s | 13.7 s |
| P50 task time after switching to Codex | 260 s | 228 s | 178 s | 108 s |
| Qwen3-14B + 8 servers, P50 task time | 288 s | 253 s | 207 s | 188 s |
AgentENV's position in this table is clear: fastest among all baselines, but behind WeEnv. This itself is favorable to AgentENV—it is the only one of the three baselines below 60 s on initialization (the other two are 111 s and 150.6 s).
This sentence must be quoted verbatim (WeEnv §9.1):
"Since our servers lack ublk kernel support, the on-demand fetching is disabled in our experiments."
ublk is the physical carrier of AgentENV's on-demand loading (the entire path in §5.3 and §5.4 is built on ublk block devices). Servers lacking ublk kernel support means AgentENV's block-level on-demand pull completely did not run—and this is precisely its core advantage over "pulling full images."
Conclusion: The 59.7 s number is not AgentENV architecture's performance, but "what AgentENV degrades to on a machine without ublk."
Immediately following in the same paragraph:
"We tune AgentENV for better performance. The file-level layers are converted into block-level ext4 layers offline and pushed to the registry, so that initialization fetches the converted layers directly and bypasses the conversion during initialization, reducing its initialization latency by 25%."
That is: moving conversion work off the initialization path, buying 25% initialization speedup.
This point is often overlooked, but it is important because it shows WeEnv's operation is not simply weakening the opponent: they both disabled AgentENV's on-demand loading (−) and moved its initialization-path conversion overhead away (+). The net bias direction is unknown—we cannot calculate from the paper whether these two offset to optimism or pessimism.
The honest conclusion is: The 5.6–14.2× multiple holds under this specific setup of "no ublk + offline pre-conversion." It is a valid measurement, but it can neither be extrapolated to "AgentENV architecture is 5.6–14.2× slower," nor used to infer DSec's performance.
WeEnv's four backends are E2B, Docker, AgentENV, WeEnv. DSec is not among them—the two papers are 6 days apart (09-19 vs 09-25), and I checked both reference lists item by item; neither cites the other.
So:
AgentENV README:29 claims "Snapshot-backed environments boot or resume in under 50 ms and pause in under 100 ms"; WeEnv measured median initialization at 59.7 s.
These two numbers differ by three orders of magnitude, seemingly contradictory, but likely measure two different things:
The real problem is not contradiction, but unalignability: AgentENV never defines the start and end points of "boot," which path it measures, under what cache state, and leaves no benchmark scripts in the repository. So these two numbers cannot be placed on the same coordinate axis. This is not unique to AgentENV—as long as a project's performance claims are not reproducible, it cannot participate in cross-institution comparison.
WeEnv's evaluation is currently the only cross-institution data; its value is giving a lower bound under restricted conditions. But when using it, four labels must be attached: AgentENV on-demand loading disabled, AgentENV offline pre-conversion speedup 25%, no DSec in the experiment, the two compared metrics have different definitions.
Setting aside "who wins," there are six conclusions in these three materials that can be directly transferred to any team doing agent infrastructure. Each is labeled with its source institution—this is also the practical use of the "three institutions" framework: they are empirical experiences from three different situations.
WeEnv's methodology can be copied directly: split RL iteration into Train / Env init / Task exec / Others four segments, normalize by total latency, then split rollout into initialization and task execution. After doing this, you know whether to invest in training, environment, or task.
WeEnv's measured results (53.4% / 39.5% initialization share, training only 15.4%–24.0%) will likely differ for other teams, but "split and measure first" itself is extremely low-cost and high-information.
Two independent instruments give same-direction but different-magnitude numbers:
| Measurement object | Access share |
|---|---|
| WeEnv (file-level layers, SWE tasks) | Median 0.81%, p99 6.02% |
| DSec (container image overall, by language) | 4.2% – 13.3% |
Both are correct because the measured granularity differs: WeEnv counts "artifact bytes actually read by the task," DSec counts "fraction of image data accessed at runtime."
Design implication: Do not copy numbers; copy the order of "measure first, then set granularity." And this magnitude difference exactly explains why both routes work—file-level layers are naturally coarser than block-level, so file-level schemes' gains depend more on "how well layers are cut," while block-level schemes' gains depend more on "how fast on-demand reads are."
Two institutions independently reached the same packaging principle:
O(m·N), O(k·N) upgrade costs to O(m), O(k); implementation cost is 30 lines of Go modifying dockerd.Two actionable corollaries:
WeEnv lists it as the third scheduling principle, with a reason weightier than any performance argument:
"A trajectory silently corrupted by starvation is worse than a lost one, as it feeds wrong signals into training."
So when OOM occurs with no tier left to raise, it chooses to terminate the task and report the cause, letting the RL framework resample, rather than letting it run sick to completion.
AgentENV gives the engineering form of the same principle at the implementation layer: failures must be graded—
| Failure type | Handling |
|---|---|
| Recoverable (capture/fork error) | Restore source sandbox to Running, source unharmed |
| Terminal | Only then require tearing down source sandbox |
| Child sandbox start failure | Must not drag down already-successful siblings (fork_sandbox_keeps_successful_siblings_when_one_start_fails) |
| Volume subvolume creation failure | cleanup_fork_volume_children() rollback |
These two things are the same principle: in the face of training data correctness, "partial success" is far more dangerous than "explicit failure."
DSec §6.4's case list should be read by everyone doing agent infrastructure. The most valuable part is not its mitigations (AppArmor + eBPF), but what the failures look like:
/bin/bash;XFS_IOC_SWAPEXT, an ioctl swapping file extent mappings to bypass—at the cost of corrupting filesystem metadata;/proc/kpagecgroup, triggering a kernel bug and crashing the kernel;yes command producing tens of GB of captured stdout.WeEnv gives an architectural answer from the other side: evaluation must run in an independent, clean environment, "so that rollout-side modifications cannot contaminate the reward computation, guarding against reward hacking."
DSec also adds an easily overlooked governance measure: builder and runtime use different accounts, and build-time residual data is cleared from the writable layer before packaging, avoiding bringing reference answers into the image.
Actionable corollary: Treat "from which unintended channels answers may leak" as an attack surface requiring continuous enumeration (internal sockets, logs, overwritten interpreters, package proxies, process filesystem), not one-time data cleaning. Two independent institutions both confirm this is a continuous adversarial process—WeEnv even writes it as a packaging cost item ("evaluator must be frequently updated against reward hacking; each change touches affected images, possibly taking days").
AgentENV's warm pool is a very clean paradigm (architecture.md:142, default.toml:338-344):
(network slot, Firecracker process) pairs;[pool.block] reusing the OverlayBD ublk device pool, shared watermarks low=2 / high=64.This and DSec's placement design are two faces of the same philosophy: DSec uses "local view overlaying its recent placements not yet reflected in periodic watcher snapshots" to eliminate coordination latency; AgentENV uses warm pool to eliminate creation latency. Neither is "making a step faster," but "moving a step off the critical path, or making it require no coordination."
The two take opposite approaches, but both write their reasons clearly, which is more important than choosing a side:
| DSec | AgentENV | |
|---|---|---|
| Approach | Explicitly designed to require no persistent state | Default in-memory, but provides optional Redis binding store + query-only replicas (services/README.md:293, services/scheduler/internal/redis_store.go:28) |
| Reason | Make apiserver / placement / watcher instances freely addable, removable, replaceable, with no expensive recovery steps; validate with periodic IaC rebuilds | Leave state authority to nodes; control plane only routes; HA needs filled by optional backend |
| Cost | Requires rebuild capability and validation mechanism (BGP + multiple instances + periodic cluster reset) | architecture.md:223's statement is too broad; the real gap is artifact indexing still only in memory under HA mode |
Common point: Neither treats the control plane as the sole authority for state. The difference is opposite paths—DSec is actively designed stateless and proves rebuild capability with periodic IaC rebuilds; AgentENV is stateful by default, then adds an optional Redis backend, and explicitly admits HA does not cover artifact indexing. The former trades constraints for elasticity; the latter trades options for elasticity.
One release read down to its config values. This one takes three releases and reads them against each other: what DeepSeek, Tencent WeChat and Moonshot each built for the same bottleneck, where their designs genuinely diverge, and which of their numbers you can actually check. The environment tax that all three measure is the same one Part 1 named and Part 2 traced through a single stack.