← All research

market

What AI expert job postings reveal about RL environment demand

3 expert marketplaces tracked daily, postings deduplicated

Heatmap of expert domains by marketplace, cell intensity showing posting volume

Expert job listings on Mercor, AfterQuery, and micro1 are a leading indicator of where labs are investing in environment development. We track active expert postings across all three platforms daily, deduplicated so cross-posted roles count once. The pattern is consistent: labs recruit domain experts months before the environments those experts build reach a training run. Demand concentrates in a few domains (coding, finance, legal, medical) while the vendor census shows 21 of 38 tracked companies sell coding environments. Where recruiting volume diverges from what vendors currently sell is where the next environment categories will appear.

Key Takeaways

  • Labs recruit domain experts on Mercor, AfterQuery, and micro1 months before the resulting environments reach training runs, which makes postings the earliest public demand signal.
  • Recruiting demand spans medical, legal, and finance, while 21 of 38 tracked vendors sell coding environments. That divergence is a map of the whitespace.
  • Reading labor data at the task level has a real research lineage, from OpenAI’s O*NET exposure scoring to Anthropic mapping millions of conversations onto occupational tasks.

What do the three marketplaces reveal?

The three platforms serve different segments of the expert-supply market:

  • Mercor operates the largest expert network, with roughly $2B in annualized gross payment volume and a 27% gross margin.
  • AfterQuery runs a YC W25 expert platform supplying frontier labs with rubrics, environments, and computer-use trajectories.
  • micro1 focuses on expert roles for AI training and evaluation.

Every active posting is classified by domain, using each platform’s native taxonomy first and falling back to title-based inference only when the platform provides no domain tag. The classification is auditable: each posting carries its source platform, native domain label, and the rule that assigned it. This matters because title-only taxonomies misfile roles: a content writer tagged multilingual because “English” appeared in a skills list, or a finance role filed under data because the title mentioned analysis.

Where does recruiting volume diverge from vendor supply?

The June vendor census shows 21 of 38 tracked companies sell coding environments, 18 sell enterprise workflows, and 18 sell computer-use environments. Math has 2 sellers. Medical has effectively one dedicated vendor. The marketplace demand signal tells a related but distinct story: labs are recruiting experts across a broader set of domains than the vendor supply covers, particularly in medical, legal, and finance roles.

Stanford’s WORKBank audit makes the same point from the worker side: it crossed 1,500 workers’ automation preferences with expert capability ratings across 844 tasks and 104 occupations, and found demand and capability rarely line up neatly (Shao et al., 2025). The gap between what vendors sell and what labs are recruiting experts to build is the whitespace the market will fill next. When a lab recruits a medical billing specialist through a marketplace, it is either building that environment internally or contracting a vendor to build it. Either way, the recruiting signal precedes the environment by months.

How reliable is marketplace data as a leading indicator?

Each posting is one observation, classified by platform and domain. Native platform taxonomy always wins over title inference, and we track the share of postings classified by native taxonomy versus inference. Postings that match no domain rule are listed as unclassified with their title and platform, so the noise is visible rather than hidden.

Task-level labor data is a legitimate instrument, not a hack. OpenAI’s GPTs are GPTs scored ONET tasks for LLM exposure and set the precedent for treating the task, not the job title, as the unit of analysis. Anthropic’s Economic Index runs the same mapping in reverse, projecting millions of real Claude conversations onto ONET tasks to see where AI is already doing the work (Handa et al., 2025). And labor economists now extract O*NET structure from job-posting corpora at the scale of 155M listings (Meisenbacher et al., 2025). Our tracker applies that same task-level lens to a narrower question: which expertise are AI labs paying to acquire right now?

The leading-indicator thesis is straightforward: a lab posting for a payroll administrator or a dental biller is not hiring for current operations. It is building an environment that requires domain expertise the lab does not have internally. The posting appears months before the environment reaches a training run, because environment construction, verifier calibration, and expert validation all precede model exposure.

What this means

Marketplace listings are the earliest public signal of where environment spend is heading. The domains where recruiting volume exceeds vendor supply are the domains where new environments will appear in the next two quarters.

FAQ

Why three marketplaces and not more?

Mercor, AfterQuery, and micro1 are the three platforms that explicitly serve AI labs seeking domain experts for training and evaluation work. Other job platforms serve broader markets and don’t isolate the AI-lab demand signal.

What is the classification noise problem?

Platform taxonomies are inconsistent. A role tagged “data” on one platform may be a finance role on another. Native taxonomy is used first and title-based inference as a fallback, with every posting’s classification rule recorded so misclassifications are auditable rather than invisible.

How far ahead of environment releases do postings appear?

The lag is months, not weeks. An expert recruited to author cases and calibrate a verifier is engaged before the environment exists, not after. By the time a benchmark launches publicly, the recruiting that built it is already six months old.