A three-part series · Agentic RL

Agentic RL Environment Engineering

The binding constraint on how capable an agent becomes is no longer the model. It is the worlds it practices in — how many, how hard, how trustworthy. This series approaches that claim from three directions: a field guide to the discipline, a close reading of one release that puts the ideas on trial, and a three-way teardown of the sandbox infrastructure three institutions actually shipped.

Apache-2.0 · every claim in all three instalments is regenerated by a gate on each push

The three instalments

Part 1 · The field guide

Your Agent Is Only as Smart as the World It Trains In

Nobody is scaling the world.

A wide-angle map of environment engineering — the discipline that decides what an agent can possibly learn. Built from roughly eighty papers, and organised around two mental models you can carry into any environment you evaluate: the GEF loop (Generate–Execute–Feedback) and the environment lifecycle (modeling → synthesis → evaluation → application).

  • Why the same 32B model swings from 6.2% to 68.2% on SWE-bench when only its training world changes.
  • The five tests a training-grade environment must pass — and which common "agent training data" quietly fails.
  • Three dirty secrets: difficulty had to become a measurement, your judge may be the weakest link, and a benchmark rots the moment you publish it.
  • The symbolic-vs-neural split, and why the honest answer is fusion: neural for throughput, code for the verdict.
  • It closes on an open call — contributing a world should be as easy as opening a pull request.
~10,000 words~80 papers & frameworks10 sectionsglossary + paper index

Read Part 1 →

Part 2 · The ground survey

How Xiaomi MiMo-V2.6 Actually Does RL

Reading the MiMo-V2.6 open stack down to the config values.

Part 1 asks whether anyone answers its open call. Part 2 examines a release that does: a 44-page report, a verl fork implementing the recipe, and the environments themselves as 3,764 runnable Docker images with verifiers and manifests attached. Every RL claim in the report is traced to the file, config value, or task environment that implements it — including three claims the released code does not implement at all.

  • Eq.(5)'s conservation is enforced numerically. Segment-level advantage reallocation ends in an assert that positive mass is preserved — because injecting net negative mass, with the shipped use_kl_loss: false and entropy_coeff: 0, flattens the policy and lengthens outputs.
  • "Reference RL framework" appears eight times in the fork, zero times in the paper — the code anonymises a system the report names in §6.4.
  • Three files carry a ByteDance copyright header and 404 upstream, on both v0.9.0 and main. Checked, not assumed.
  • Four decoupled clip bounds are described; only two values ship. Across the whole repository the only values present are 0.2 and 0.28.
  • The data-to-environment join is exact. 3,764 mapping rows, 3,764 Docker Hub tags; code and cyber are one image per task, precisely.
  • One GRPO group never spans two harnesses, and the anti-hack guard ships default-off, stage 1 only.
13 sections735 Python files read7,780 task rows3,764 images

Read Part 2 →

Part 3 · The three-way teardown

How DSec, WeEnv and AgentENV Actually Engineer Sandboxes

Three institutions, three released systems, one shared diagnosis.

Parts 1 and 2 each looked at one thing in depth. Part 3 puts three side by side — DeepSeek's DSec, Tencent WeChat's WeEnv and Moonshot's AgentENV — and reads them against the same measurements. All three independently arrived at the same diagnosis: the environment tax, where initialisation eats up to 53.4% of iteration time.

  • One platform, four isolation backends. DSec's answer to workloads that cannot share a runtime: FnCall, container, Firecracker microVM, QEMU VM.
  • The full lifecycle, governed as a supply chain. WeEnv takes packaging, initialisation and provisioning end to end — and cuts init from 53.4% to 9.1%.
  • Storage as the substrate. AgentENV makes every layer of state, memory included, a block device that can be snapshotted, forked and paused.
  • Three orthogonal bets. Density, lifecycle, startup latency — the three draw the boundary of "environment" in completely different places.
  • Claims are separated from measurements. The one cross-institution benchmark carries four qualifications, including that the system under test had its core mechanism disabled.
2 languages9 sections3 papers & codebases35 tables

Read Part 3 →  ·  中文版

How the claims are checked

Both instalments make countable claims, so the repository carries the checks that would catch them drifting. tools/check_claims.py is a standard-library gate over link resolution, the count ledger, and offline-renderability of every served page. tools/test_check_claims.py is what makes that gate trustworthy: it builds throwaway repositories, injects one deliberate defect into each, and asserts the gate fails — 21 mutations, all detected. Without it, a gate that silently stopped checking anything would look exactly like a gate that works.

The raw commands and their actual output behind every finding live in evidence/VERIFICATION.md, and the machine-readable counts in evidence/offsets.json.