← All research

craft

Which jobs can become RL environments?

A physical-action veto blocks construction, not just the score

Funnel of candidate occupations passing gate checkpoints; rejected ones drop at a marked veto gate

A screening framework answers the question every vendor and lab faces: can this job be simulated with a verifiable reward, or not? The framework evaluates candidate occupations across ten dimensions, from computer-work share to output verifiability, with one hard rule: a critical physical dependency vetoes construction outright, no score adjustment. A job that requires hands-on equipment operation, physical inventory observation, or in-person client interaction can’t become a sandbox environment, no matter how well it scores everywhere else. The framework explains the gap between the 99 catalogued environments and the 66 tracked occupations that have none.

Key Takeaways

  • Ten dimensions decide whether a job can become an RL environment, and nine of them are negotiable. A critical physical dependency is not: it vetoes construction outright.
  • Three calibration workflows pass the gate with computer-work shares of 0.82 to 0.96, and each names exactly which physical capabilities it excludes.
  • The framework predicts the market: Surge’s Chartography passes every screen, while construction trades, dental work, and farming stay unbuildable at any budget.

What are the ten screening dimensions?

Each candidate occupation is evaluated on ten dimensions:

  • workflow modality
  • computer-work share
  • sandbox-action coverage
  • digital-evidence availability
  • action observability
  • output verifiability
  • collaboration simulability
  • native-tool fidelity
  • long-horizon depth
  • critical physical dependency

The first nine are scored on a 0–1 scale. The tenth is a binary veto, and it’s the load-bearing rule. A job whose core competency requires physical manipulation of equipment, hands-on bench operation, or in-person client interaction can’t be simulated at a desk. No amount of computer-work share compensates. The framework rejects the job before scoring the other dimensions, because building a sandbox for a physical job measures the wrong thing: tool adaptation, not competence.

The task-level lens has a strong precedent. OpenAI’s GPTs are GPTs scored ONET tasks for LLM exposure and drew the same hard line from the other direction: tasks requiring physical action scored zero exposure by rubric definition. And Anthropic’s mapping of millions of real conversations onto ONET tasks shows usage concentrating in exactly the computer-mediated task families the framework passes (Handa et al., 2025).

What do the calibration points show?

Three occupational workflows screened with the framework illustrate it in practice:

  • Failure diagnosis (remote reliability triage and maintenance work-package authoring) scores 0.82 on computer-work share. The scoped workflow excludes physical inspection, lockout execution, and repair workmanship. The desk-based triage is a real competency.
  • Engineering test analysis scores 0.92. The scoped workflow excludes hands-on bench operation and sensor installation. The data audit, design critique, and decision-making is strongly computer-mediated.
  • Audit workpaper completion scores 0.96. Audit is computer-native: source records, formulas, evidence links, and conclusions are all digital.

All three pass the screening gate. None has a critical physical dependency for the scoped workflow. The excluded capabilities are explicitly named, not hidden. A screening framework that can’t state what it does not measure is marketing, not engineering.

The existence proofs are piling up on the pass side of the gate. SWE-Lancer converted 1,400+ real freelance software jobs, with real dollar values attached, into end-to-end verifiable tasks. GDPval built expert-graded tasks from real work products across 44 occupations. Both are demonstrations that when work is fully digital and its output is checkable, someone will turn it into an environment.

What does the week’s news confirm?

Surge launched Chartography on July 16: expert-graded professional chart reading. The benchmark passes every screen in the framework. Fully computer-mediated, deterministic evidence, verifiable output. A chart-reading task requires no physical manipulation, produces a digital artifact, and can be graded against expert rubrics. It’s the kind of work the framework predicts is environment-eligible.

The contrast is instructive. The 66 tracked occupations without environments include construction trades, dental work, and farming, jobs with critical physical dependencies that veto sandbox construction. They also include accounting, legal work, and consulting, jobs that score high on computer-work share and output verifiability but have no environment yet. The framework explains both the empty cells and the filled ones.

What this means

The screening framework is the first filter every environment builder applies, whether explicitly or implicitly. The occupations that pass are the addressable market. The occupations that fail at the physical-dependency veto are not, and no amount of vendor effort will change that.

FAQ

What is a critical physical dependency?

A job has a critical physical dependency when its core competency requires manipulating physical objects, equipment, inventory, patients, materials, in ways that cannot be represented at a desk. The framework vetoes these jobs outright because a sandbox simulation would measure adaptation to substitute tools, not the competence the job actually requires.

Why is OS fidelity a measured variable?

An environment that grades audit work on LibreOffice measures tool adaptation, not audit competence. Windows and native Office are selected for audit work because workbook formulas, review behavior, and final artifacts are part of the target skill. The framework scores native-tool fidelity as a dimension, not a binary, because the question is whether the tool choice affects what is being measured.

What is Chartography?

Chartography is a Surge AI benchmark launched July 16 for expert-graded professional chart reading. It passes every screen in the framework: fully computer-mediated, deterministic evidence, verifiable output.