A three-part series · Agentic RL
The binding constraint on how capable an agent becomes is no longer the model. It is the worlds it practices in — how many, how hard, how trustworthy. This series approaches that claim from three directions: a field guide to the discipline, a close reading of one release that puts the ideas on trial, and a three-way teardown of the sandbox infrastructure three institutions actually shipped.
Nobody is scaling the world.
A wide-angle map of environment engineering — the discipline that decides what an agent can possibly learn. Built from roughly eighty papers, and organised around two mental models you can carry into any environment you evaluate: the GEF loop (Generate–Execute–Feedback) and the environment lifecycle (modeling → synthesis → evaluation → application).
Reading the MiMo-V2.6 open stack down to the config values.
Part 1 asks whether anyone answers its open call. Part 2 examines a release that
does: a 44-page report, a verl fork implementing the recipe, and the
environments themselves as 3,764 runnable Docker images with verifiers
and manifests attached. Every RL claim in the report is traced to the file, config
value, or task environment that implements it — including three claims the
released code does not implement at all.
assert that positive mass is preserved — because injecting net negative mass, with the shipped use_kl_loss: false and entropy_coeff: 0, flattens the policy and lengthens outputs.v0.9.0 and main. Checked, not assumed.0.2 and 0.28.Three institutions, three released systems, one shared diagnosis.
Parts 1 and 2 each looked at one thing in depth. Part 3 puts three side by side — DeepSeek's DSec, Tencent WeChat's WeEnv and Moonshot's AgentENV — and reads them against the same measurements. All three independently arrived at the same diagnosis: the environment tax, where initialisation eats up to 53.4% of iteration time.
Both instalments make countable claims, so the repository carries the checks that
would catch them drifting. tools/check_claims.py is a standard-library
gate over link resolution, the count ledger, and offline-renderability of every
served page. tools/test_check_claims.py is what makes that gate
trustworthy: it builds throwaway repositories, injects one deliberate defect into
each, and asserts the gate fails — 21 mutations, all detected.
Without it, a gate that silently stopped checking anything would look exactly like
a gate that works.
The raw commands and their actual output behind every finding live in evidence/VERIFICATION.md, and the machine-readable counts in evidence/offsets.json.