Research & analysis
Original data and analysis on the RL environment economy — vendors, benchmarks, environment craft, and the infrastructure underneath.

Who defines ground truth for AI agents?
The quality bar separating environments labs trust from environments labs return: separated roles, published calibration thresholds, and models that never approve truth.

Running AI agents in microVMs with zero inbound ports
Long-horizon episodes need real operating systems, real applications, and hard isolation; the emerging pattern is microVMs with zero inbound ports.

How AI agents cheat their training environments
Under RL pressure agents game environments in predictable ways, from format masquerading as competence to quiet memorization of leaked test structure. The counters are becoming standard.

Training AI on office work made it better at coding
Environment training transfers across domains. A model post-trained on non-coding office tasks gained 5.8 points on SWE-Bench Pro, which changes what labs are buying.

What Mercor's APEX-Accounting benchmark measures, and what it misses
Mercor and Ramp's APEX-Accounting is a serious benchmark for reconciliations and closing the books. But a benchmark cannot carry evidence integrity, which is what audit work is actually graded on.

What a production-grade RL environment spec looks like
Most tracked environments publish no verifier calibration and no frozen corpus; the ones labs pay for fix state machine, reward, and splits in a contract before any code.

Which jobs can become RL environments?
A screening framework answers the question every vendor and lab faces: can this job be simulated with a verifiable reward, or not?

Why Mercor bought Deeptune
Mercor's Deeptune acquisition says the constraint has shifted from expert networks to the environments themselves. The vendor data shows the whitespace the deal targets.

3,029 AI benchmarks, 99 RL environments: the field is over-indexed on evaluation
The benchmark pile keeps growing while catalogued RL environments stay rare and static. Expert verification, not benchmark count, moves the numbers.

AI benchmarks cover only 3.5% of real work
Joining 202 O*NET occupational tasks against 223 benchmarks with an LLM judge yields definitive coverage of 3.5%. Benchmark success overstates workflow competence.

What AI expert job postings reveal about RL environment demand
Marketplace listings are a leading indicator: labs recruit experts months before the environments those experts build reach training runs.

The RL environment industry is 38 companies, mostly under 50 people
A cottage industry of mostly sub-50-person vendors supplies the most capitalized labs on earth.

The economics of selling RL environments to AI labs
Frontier labs are shifting spend from labeled data to executable environments, and vendor margins show the leverage sits in verification, not labor.