Research & analysis

Original data and analysis on the RL environment economy — vendors, benchmarks, environment craft, and the infrastructure underneath.

findings

Who defines ground truth for AI agents?

The quality bar separating environments labs trust from environments labs return: separated roles, published calibration thresholds, and models that never approve truth.

infrastructure

Running AI agents in microVMs with zero inbound ports

Long-horizon episodes need real operating systems, real applications, and hard isolation; the emerging pattern is microVMs with zero inbound ports.

craft

How AI agents cheat their training environments

Under RL pressure agents game environments in predictable ways, from format masquerading as competence to quiet memorization of leaked test structure. The counters are becoming standard.

findings

Training AI on office work made it better at coding

Environment training transfers across domains. A model post-trained on non-coding office tasks gained 5.8 points on SWE-Bench Pro, which changes what labs are buying.

findings

What Mercor's APEX-Accounting benchmark measures, and what it misses

Mercor and Ramp's APEX-Accounting is a serious benchmark for reconciliations and closing the books. But a benchmark cannot carry evidence integrity, which is what audit work is actually graded on.

craft

What a production-grade RL environment spec looks like

Most tracked environments publish no verifier calibration and no frozen corpus; the ones labs pay for fix state machine, reward, and splits in a contract before any code.

craft

Which jobs can become RL environments?

A screening framework answers the question every vendor and lab faces: can this job be simulated with a verifiable reward, or not?

market

Why Mercor bought Deeptune

Mercor's Deeptune acquisition says the constraint has shifted from expert networks to the environments themselves. The vendor data shows the whitespace the deal targets.

findings

3,029 AI benchmarks, 99 RL environments: the field is over-indexed on evaluation

The benchmark pile keeps growing while catalogued RL environments stay rare and static. Expert verification, not benchmark count, moves the numbers.

findings

AI benchmarks cover only 3.5% of real work

Joining 202 O*NET occupational tasks against 223 benchmarks with an LLM judge yields definitive coverage of 3.5%. Benchmark success overstates workflow competence.

market

What AI expert job postings reveal about RL environment demand

Marketplace listings are a leading indicator: labs recruit experts months before the environments those experts build reach training runs.

market

The RL environment industry is 38 companies, mostly under 50 people

A cottage industry of mostly sub-50-person vendors supplies the most capitalized labs on earth.

market

The economics of selling RL environments to AI labs

Frontier labs are shifting spend from labeled data to executable environments, and vendor margins show the leverage sits in verification, not labor.