The RL environment industry is 38 companies, mostly under 50 people
31 of 38 tracked vendors have ≤50 employees

Our census of 38 tracked RL environment vendors finds 31 with 50 or fewer employees. The companies supplying the most capitalized labs on earth are a cottage industry. Scale AI and Turing are the only firms above 200 staff. Everyone else, including the vendors whose benchmarks appear in frontier system cards, runs on a small team. And small doesn’t mean marginal: this same week, AfterQuery, a 51–200-person YC W25 lab, published a +21.4% net win-loss gain on GDPval from on-policy distillation. Small shops move frontier-model numbers.
Key Takeaways
- 31 of 38 tracked RL environment vendors have 50 or fewer employees; only Scale AI and Turing top 200.
- AfterQuery moved GDPval, OpenAI’s benchmark of economically valuable work, by +21.4% net win-loss with on-policy distillation. Headcount didn’t do that; expert signal did.
- The concentration risk cuts both ways: one lost lab contract can erase a vendor, and one acquisition can reset the market.
What does the headcount distribution look like?
The modal band is 11–50 employees, holding 19 of 38 vendors. That’s the size where a company can field a few domain teams, build a benchmark, and serve one to three lab customers. It’s also the size where a single acquisition can reset the competitive landscape.
The two firms above 200 are Scale AI and Turing. Both are legacy data-labeling companies that expanded into environments, not environment-native startups.
Who are the sub-50-person vendors?
Three vendors illustrate the cottage-industry pattern:
- Mechanize, founded April 2025 by ex-Epoch AI researchers. High-fidelity coding environments, 11–50 people, a $9.1M seed at a $500M post.
- Fleet AI, 11–50 people. Enterprise gyms replicating Salesforce and Excel for frontier labs.
- Gray Swan AI, a Carnegie Mellon spinout in Pittsburgh, 11–50 staff. Adversarial red-teaming and runtime protection.
AfterQuery sits at 51–200, just above the sub-50 band. The YC W25 lab publishes Terminal-Bench, FinanceQA, and IDE-Bench, and supplies expert-generated rubrics and environments to frontier labs. Its size is the exception, not the rule.
How do small vendors prove value?
AfterQuery published a +21.4% net win-loss gain on GDPval from on-policy distillation this same week. A sub-200-person shop moved a frontier evaluation metric by double digits. The method itself is public research: on-policy distillation trains the student on its own outputs with teacher feedback, the GKD recipe from Agarwal et al. (2023). What the vendor adds is the expert-generated signal the method consumes.
That’s the mechanism by which cottage-industry vendors earn lab contracts: not headcount or revenue, but a measured uplift on a benchmark a lab already tracks. The academic record says this is exactly how it should work. LIMO elicited sophisticated reasoning from a few hundred curated examples, and Stanford’s s1 matched much larger training runs with a curated set of 1,000. Data quality beats data volume, and quality is a thing a 20-person expert shop can win on.
The pattern repeats across the vendor list. The companies winning lab spend can point to a specific number: a benchmark result, a training uplift, a system-card citation. The size of the engineering org doesn’t enter into it.
What this means
The industry’s concentration risk is real. 31 of 38 vendors are small enough that a single lost lab contract or a single acquisition removes them from the market. For acquirers, the sub-50 band is where talent and tooling are cheapest to buy. For labs, the risk is that a vendor they depend on is one failed quarter from disappearing.
FAQ
Why are most RL environment vendors so small?
The market has 4–6 buyers. Frontier labs are the primary customers, and each lab runs a named internal environments team that supplements rather than outsources. A vendor needs enough staff to build environments and grade them, but the demand doesn’t support scaling beyond a handful of lab contracts.
What is GDPval?
GDPval is OpenAI’s benchmark for professional work tasks across real occupations, built from expert-authored tasks spanning 44 occupations in the top GDP sectors (Patwardhan et al., 2025). Vendors publish training and evaluation results on it to demonstrate that their environments or data improve frontier-model performance on occupational work.
Does the headcount distribution predict acquisition targets?
The sub-50 band is where acquirers find talent and tooling at the lowest price. The 200+ firms are too expensive for most buyers and too entrenched to sell. When consolidation comes to this market, it will come through the small end.