The economics of selling RL environments to AI labs
27% gross margin on ~$2B annualized gross (Mercor)

Frontier labs are shifting spend from labeled data to executable environments, and the vendor margins tell you where the leverage sits. Mercor runs roughly $2B in annualized gross payment volume at a 27% gross margin and is valued at $10B after a $350M Series C. AfterQuery closed a $30M Series A at a $300M valuation and says it has since passed a $100M revenue run rate. Neither margin profile looks anything like software. The category winner sells expert labor organized around verification, and the labs buying it now cite vendor evals in their own system cards. The money is moving from annotation budgets to environment budgets, and it’s flowing to the companies that can grade what a model does, not just label what it sees.
Key Takeaways
- Mercor processes ~$2B in annualized gross payment volume at a 27% gross margin, and the market prices that services business like infrastructure: $10B post-money.
- AfterQuery raised a $30M Series A at a $300M valuation and reports crossing a $100M revenue run rate on expert-generated rubrics, environments, and trajectories.
- Labs put vendor evals in system cards this month (Anthropic cited Surge’s GDP.pdf and Riemann-bench on June 7), which is what verification-as-a-product looks like.
What do the margins say about the business?
Mercor’s 27% gross margin is a services margin, not a software margin. The company processes roughly $2B in annualized gross payment volume, money flowing from labs to contracted experts, and keeps about a quarter after expert pay and operational costs. A $10B valuation on that base prices in one belief: environment demand compounds with every model release. It doesn’t price in a cost structure that scales sublinearly, because there isn’t one. The economics of paying experts are stubborn, and the research on annotation labor says quality tracks pay and instruction clarity, not headcount (Laux et al., 2023).
AfterQuery’s shape matches. A $30M Series A at a $300M valuation, a self-reported $100M run rate a few months later, and a product line of expert-generated rubrics, agent environments, and computer-use trajectories sold to frontier labs. Both companies are priced as infrastructure for a recurring demand cycle, not as one-off data vendors.
What did Mercor’s thesis essay say?
Mercor published “The Economy will Become an RL Environment Machine” in September 2025, arguing that the market for humans teaching models is sized by what models cannot yet do. As capability advances, the work shifts from labeling correct answers to building the environments in which correct behavior is defined. The essay said the quiet part: the addressable market is the gap between model competence and occupational competence, and that gap is where verification becomes the product.
The training research backs the thesis. RLVR, reinforcement learning with verifiable rewards, swaps learned reward models for deterministic checkers (Lambert et al., 2024), and DeepSeek-R1 showed frontier reasoning emerging from RL against verifiable signals rather than piles of human-labeled data. If the trainable signal is verification, the scarce input is whoever can define and grade correctness. The margin data is consistent: if the product were labeled data, margins would compress as labeling commoditizes. If the product is verification, the moat deepens as tasks get harder.
Are vendors becoming frontier infrastructure?
In the first week of June, two labs leaned on vendor evals in public:
- Anthropic cited Surge AI’s GDP.pdf and Riemann-bench in the Fable 5 system card (June 7).
- Microsoft used Surge human evals to benchmark MAI-Thinking-1 (June 2).
Vendor benchmarks are appearing in lab release notes, not just in vendor marketing.
That’s the structural shift. When a lab cites a vendor’s eval in a system card, the vendor is supplying part of the frontier’s measurement layer. The spend that used to go to annotation pipelines now goes to companies that can build and grade executable environments, and labs keep raising the stakes: Meta’s ScaleRL study burned 400,000 GPU-hours just to establish how RL post-training compute scales (Khatri et al., 2025). Every scaling curve like that is a purchase order for more environments. The 27% margin says the work is still labor-intensive. The system-card citations say the output is becoming indispensable.
What this means
The economics point one direction: verification, not labor supply, is where margin accrues. Labs buying environments are buying the ability to distinguish competence from performance, and the vendors that can do that at scale are the ones capturing the spend.
FAQ
Why is 27% gross margin a services margin?
Software businesses typically run 70-80% gross margins because serving one more customer costs almost nothing. Mercor’s 27% reflects the cost of paying experts for each unit of work. The marginal cost doesn’t collapse, and that’s the signature of a services business.
What is GDP.pdf?
GDP.pdf is a Surge AI benchmark for professional document tasks. Frontier labs cite it in system cards as an external measure of professional-work competence, which is why it shows up in Anthropic’s Fable 5 release and OpenAI’s GPT-5.6 materials.
If margins are services-level, why are valuations software-level?
Because every model release manufactures demand for the next batch of environments. The revenue is recurring even though the contracts are not, and buyers describe exactly this cycle: a new frontier model ships, the post-training position resets, and a fresh order follows.