Mechanize
Mechanize is a small, elite San Francisco vendor (founded April 2025 by ex-Epoch AI researchers Matthew Barnett, Tamay Besiroglu, and Ege Erdil) that builds a …
A managed leaderboard of RL environments, benchmarks, and model scores — scraped, normalized, and synthesized by agents from public research sources.
The leaderboard now tracks 535 benchmarks across 98 environments and companies, with 291 models and 9 tools. Notable new benchmarks include GBA Eval for coding Game Boy Advance emulators, VADER for vulnerability assessment, and FinanceQA for financial analysis in LLMs. Key companies featured are Mechanize, AfterQuery, and Bespoke Labs, with tools like Bespoke Curator for synthetic data curation and Evalchemy for evaluation workflows gaining attention. The OpenThoughts dataset collaboration with DataComp highlights ongoing advances in open reasoning benchmarks.
Loading from /api/site…
Top BenchLM model scores from the latest scrape. Scores parsed from public leaderboard data.
| Rank | Model | Score | Creator | |
|---|---|---|---|---|
| 1 | 83.21 | Anthropic | ||
| 2 | 83.07 | Anthropic | ||
| 3 | 82.96 | Anthropic | ||
| 4 | 82.00 | OpenAI | ||
| 5 | 80.50 | Moonshot AI | ||
| 6 | 79.91 | Alibaba | ||
| 7 | 76.58 | Anthropic | ||
| 8 | 75.54 | |||
| 9 | 75.42 | xAI | ||
| 10 | 73.37 | OpenAI | ||
| 11 | 73.36 | OpenAI | ||
| 12 | 71.87 | Anthropic | ||
| 13 | 71.60 | Alibaba | ||
| 14 | 68.74 | MiniMax | ||
| 15 | 68.02 | Tencent | ||
| 16 | 67.98 | Anthropic | ||
| 17 | 67.37 | |||
| 18 | 67.04 | Thinking Machines Lab | ||
| 19 | 66.94 | Xiaomi | ||
| 20 | 66.91 | Z.AI | ||
| 21 | 66.02 | Alibaba | ||
| 22 | 65.93 | OpenAI | ||
| 23 | 65.61 | Z.AI | ||
| 24 | 65.29 | |||
| 25 | 64.79 | Alibaba | ||
| 26 | 64.78 | Anthropic | ||
| 27 | 64.67 | |||
| 28 | 64.53 | Anthropic | ||
| 29 | 63.95 | xAI | ||
| 30 | 63.63 | Anthropic | ||
| 31 | 63.41 | xAI | ||
| 32 | 63.34 | Z.AI | ||
| 33 | 63.06 | MiniMax | ||
| 34 | 62.36 | Xiaomi | ||
| 35 | 61.67 | Meta | ||
| 36 | 61.36 | |||
| 37 | 61.16 | DeepSeek | ||
| 38 | 60.79 | InternScience | ||
| 39 | 60.66 | Z.AI | ||
| 40 | 60.43 | xAI | ||
| 41 | 60.21 | Moonshot AI | ||
| 42 | 60.15 | |||
| 43 | 59.99 | Alibaba | ||
| 44 | 59.82 | Alibaba | ||
| 45 | 59.74 | |||
| 46 | 59.64 | Alibaba | ||
| 47 | 59.61 | xAI | ||
| 48 | 59.36 | OpenAI | ||
| 49 | 59.16 | Xiaomi | ||
| 50 | 59.08 | OpenAI |
RL environment vendors and research orgs tracked across sources.
Mechanize is a small, elite San Francisco vendor (founded April 2025 by ex-Epoch AI researchers Matthew Barnett, Tamay Besiroglu, and Ege Erdil) that builds a …
AfterQuery is a San Francisco applied-research lab and data platform (YC W25) that supplies frontier AI labs with expert-generated human data (SFT, RL rubrics), …
Bespoke Labs is an applied AI research lab (Mountain View, CA, founded 2024) focused on data curation and RL-environment curation for training and evaluating …
Huzzle Labs is the AI division of London-based talent platform Huzzle (founded ~2020 by Ingmar Klein, Parham Rakhshanfar, and Amit Choudhary). It positions …
Fleet AI builds high-fidelity reinforcement-learning training environments ('gyms') that replicate enterprise software such as Salesforce and Excel, plus …
Datacurve is a YC W24 commercial data vendor that supplies expert-curated frontier coding data, RLHF traces, and repository-wide reinforcement learning …
Proximal is a San Francisco-based (with a Bangalore presence) research lab for coding data, building high-fidelity, long-horizon reinforcement learning …
Gray Swan AI is a Pittsburgh-based AI security company spun out of Carnegie Mellon, offering adversarial red-teaming and runtime protection for AI models and …
Veris AI sells a high-fidelity simulation platform plus a production runtime that let enterprises train, evaluate, and govern AI agents against mocked …
Chakra Labs runs Dojo, an open/collaborative reinforcement-learning environment hub for computer-use agents, offering deterministic, frame-accurate clones of …
Andon Labs is a Y Combinator-backed (W24) startup, formerly Vectorview, building benchmarks and evaluations for AI agents' long-horizon coherence and safety …
HUD (YC W25, formerly hud.so) is a platform for building reinforcement-learning environments and evaluations for computer-use and browser agents. It lets teams …
Vals AI is an independent, third-party benchmarking and evaluation platform that scores LLMs and AI applications (copilots, RAG, agents) on rigorous, …
Halluminate (YC S25, founded 2024, San Francisco) builds managed reinforcement-learning sandbox environments, simulated applications, and human/annotation data …
Matrices builds reinforcement-learning training environments for frontier AI labs to train agents that use computers and browsers like humans, described as a …
BenchFlow is an early-stage, YC-backed open-source 'environment lab' building evaluation infrastructure and a community Benchmark Hub for AI agents, with …
Collinear AI operates a 'Simulation Lab' (SimLab) that builds sandboxed, stateful RL environments simulating enterprise users, tools (Jira, ServiceNow, Shopify, …
Refresh (YC X25) builds simulation engines / RL environments with verifiable rewards for coding and computer use, partnering with frontier labs and enterprises …
Vmax is a San Francisco reinforcement-learning startup (founded 2025 by three RL/robotics PhDs from UCL and UPenn) that automates the conversion of proprietary …
Andromede is an early-stage RL data lab that programmatically generates RL environments, tasks, and verifiers from real-world data for post-training and …
Plato (plato.so, Plato Technologies, Inc.) builds simulated worlds for training and evaluating browser and computer-use agents, recreating real …
AIChamp builds custom reinforcement-learning environments and 'Virtual Gym' simulations for training and evaluating tool-using AI agents on long-horizon, …
Habitat Inc is an early-stage commercial vendor (2-10 employees, New York HQ) building reinforcement-learning environments for white-collar / work automation, …
Scale AI is the data-labeling and AI-data incumbent that has extended into RL environments, offering simulated web apps, macOS/Windows-like desktop VMs, and …
Modal (Modal Labs) is a New York-based, Python-native serverless cloud purpose-built for AI/ML workloads, providing on-demand GPU/CPU compute, fast-booting …
Mercor is a venture-backed expert-marketplace and AI-training-data company that organizes a network of ~30,000+ domain experts (doctors, lawyers, bankers, …
Surge AI is a bootstrapped, high-revenue human-data and RLHF labeling leader serving frontier AI labs, which has expanded into agentic RL environments via its …
Prime Intellect operates an open-source RL stack - the Environments Hub (2,500+ community RL environments), the Verifiers library and prime-rl training …
Daytona provides secure, elastic, programmatic sandboxes ('computers') that AI agents and developers can spin up in under ~90ms to run untrusted AI-generated …
Deeptune was a New York-based startup building managed reinforcement-learning environments ('training gyms') for computer-use and code, where AI agents practice …
E2B provides open-source, secure cloud sandboxes (built on Firecracker microVMs) for running AI-generated code and AI agents, offered as a hosted API with …
Runloop sells cloud-hosted, isolated micro-VM 'devboxes' plus blueprints, snapshots and benchmark/eval tooling that give AI coding agents a secure execution …
General Reasoning is an AI research lab (operating research hub in London; legal entity General Reasoning, Inc. registered in the US) building open RL …
Cua (trycua, YC X25) is open-source MIT-licensed infrastructure for computer-use agents, providing cloud and self-hosted sandboxes across macOS, Windows, Linux, …
Sepal AI was a YC-backed (S24) San Francisco data-research company that built high-quality training data, expert-graded evaluation benchmarks, and …
Good Start Labs is a 2025 Every spin-out that builds game-based environments to generate reinforcement-learning data and evaluate frontier models, using both …
Morph (Morph Labs) provides snapshot-based VM compute for AI agents via its Infinibranch / Liquid Metal technology, which can snapshot, branch, and restore …
Turing is a large AGI-infrastructure and engineering-services company that supplies frontier AI labs with coding data, human expertise, and RL/evaluation data …
Thiel Fellowship, Harvard, Georgetown
Palantir, Meta, Scale AI
Stanford, Coinbase, Y Combinator
Stanford, Rover, Yahoo
Google, Meta, Twitter
Stanford, UW, Meta
Stanford, Berkeley
Wharton, Penn, UBC
Epoch AI
Meta AI, Google, Amazon
Waterloo
Capital One, Meta, Cornell
Stanford, Scale, Turing
Prime Intellect, Cursor
Lazard, McKinsey, EF
Google DeepMind, UC Berkeley
Hugging Face, Salesforce, Stanford
Berkeley, Asimov, Y Combinator
Elastic, Sumo Logic, Dynatrace
Turing, Aera Technology, Oracle
Terminal, Red Hat
Microsoft, Twitter
G-Research, Etched, ETH Zurich
AWS, DeepMind, Prime Intellect
Meta, Conjecture, Aleph Alpha
Y Combinator, IIT Madras
Waymo, Nuro, Netflix
Penn M&T, Wharton
MosaicML, MSR, EF
Oxford
UCL
Hebbia
Mercor, Anthropic, MSL
Turing
Exa, Palantir, Mercado Libre
HRT, Jane Street, Mercor
Akamai
Microsoft, L/S, Snap
OpenAI, Google Brain, EleutherAI
METR, AWS, Hume
Meta, Microsoft, Cornell
CMU, MSR, RSAC Labs
Public benchmarks and eval suites referenced by tracked companies.
Mechanize · Coding, Private Codebases
AfterQuery · Coding, Computer Use, Enterprise Workflows
AfterQuery · Coding, Computer Use, Enterprise Workflows
AfterQuery · Coding, Computer Use, Enterprise Workflows
AfterQuery · Coding, Computer Use, Enterprise Workflows
Bespoke Labs · Long-Horizon
Bespoke Labs · Long-Horizon
Bespoke Labs · Long-Horizon
Bespoke Labs · Long-Horizon
Bespoke Labs · Long-Horizon
Huzzle Labs · Coding, Computer Use, Enterprise Workflows, Long-Horizon, Private Codebases
Fleet AI · Computer Use, Enterprise Workflows
Datacurve · Coding, Private Codebases
Proximal · Coding, Long-Horizon, Private Codebases
Gray Swan AI
Gray Swan AI
Gray Swan AI
Veris AI · Enterprise Workflows
Chakra Labs · Computer Use
Chakra Labs · Computer Use
Chakra Labs · Computer Use
Andon Labs · Computer Use, Long-Horizon
Andon Labs · Computer Use, Long-Horizon
Andon Labs · Computer Use, Long-Horizon
Andon Labs · Computer Use, Long-Horizon
Andon Labs · Computer Use, Long-Horizon
Andon Labs · Computer Use, Long-Horizon
HUD · Computer Use, Enterprise Workflows
HUD · Computer Use, Enterprise Workflows
Vals AI
Vals AI
Vals AI
Vals AI
Halluminate · Computer Use, Enterprise Workflows
Halluminate · Computer Use, Enterprise Workflows
Halluminate · Computer Use, Enterprise Workflows
BenchFlow · Coding, Computer Use, Enterprise Workflows
BenchFlow · Coding, Computer Use, Enterprise Workflows
BenchFlow · Coding, Computer Use, Enterprise Workflows
Vmax · Coding, Long-Horizon
Vmax · Coding, Long-Horizon
Scale AI · Coding, Computer Use, Enterprise Workflows, Long-Horizon
Mercor · Coding, Enterprise Workflows, Long-Horizon
Mercor · Coding, Enterprise Workflows, Long-Horizon
Mercor · Coding, Enterprise Workflows, Long-Horizon
Mercor · Coding, Enterprise Workflows, Long-Horizon
Surge AI · Enterprise Workflows, Long-Horizon
Surge AI · Enterprise Workflows, Long-Horizon
Surge AI · Enterprise Workflows, Long-Horizon
Surge AI · Enterprise Workflows, Long-Horizon
Prime Intellect · Coding, Enterprise Workflows, Long-Horizon, Math
Prime Intellect · Coding, Enterprise Workflows, Long-Horizon, Math
Deeptune · Coding, Computer Use, Enterprise Workflows, Private Codebases
Runloop · Coding
General Reasoning · Coding, Long-Horizon
General Reasoning · Coding, Long-Horizon
General Reasoning · Coding, Long-Horizon
Cua · Computer Use
Sepal AI · Enterprise Workflows, Long-Horizon, Math
Good Start Labs · Long-Horizon
Good Start Labs · Long-Horizon
Good Start Labs · Long-Horizon
Good Start Labs · Long-Horizon
Morph · Coding
Morph · Coding
Morph · Coding
Morph · Coding
Mercor · Multi-Domain
Handshake · Multi-Domain
Scale · Multi-Domain, Code
Turing · Code
Surge · Multi-Domain
Snorkel · Multi-Domain, Code
micro1 · Multi-Domain
AfterQuery · Multi-Domain
Mechanize · Code
Patronus AI · Multi-Domain, Code
Datacurve · Code
Halluminate · Finance, Enterprise
Pareto · Multi-Domain, Tool Use
RL environments and tooling surfaces discovered across the ecosystem.
Mechanize · Mechanize is a small, elite San Francisco vendor (founded April 2025 by ex-Epoch AI researchers Matthew Barnett, Tamay …
AfterQuery · AfterQuery is a San Francisco applied-research lab and data platform (YC W25) that supplies frontier AI labs with …
Bespoke Labs · Bespoke Labs is an applied AI research lab (Mountain View, CA, founded 2024) focused on data curation and RL-environment …
Huzzle Labs · Huzzle Labs is the AI division of London-based talent platform Huzzle (founded ~2020 by Ingmar Klein, Parham …
Fleet AI · Fleet AI builds high-fidelity reinforcement-learning training environments ('gyms') that replicate enterprise software …
Datacurve · Datacurve is a YC W24 commercial data vendor that supplies expert-curated frontier coding data, RLHF traces, and …
Proximal · Proximal is a San Francisco-based (with a Bangalore presence) research lab for coding data, building high-fidelity, …
Gray Swan AI · Gray Swan AI is a Pittsburgh-based AI security company spun out of Carnegie Mellon, offering adversarial red-teaming and …
Veris AI · Veris AI sells a high-fidelity simulation platform plus a production runtime that let enterprises train, evaluate, and …
Chakra Labs · Chakra Labs runs Dojo, an open/collaborative reinforcement-learning environment hub for computer-use agents, offering …
Andon Labs · Andon Labs is a Y Combinator-backed (W24) startup, formerly Vectorview, building benchmarks and evaluations for AI …
HUD · HUD (YC W25, formerly hud.so) is a platform for building reinforcement-learning environments and evaluations for …
Vals AI · Vals AI is an independent, third-party benchmarking and evaluation platform that scores LLMs and AI applications …
Halluminate · Halluminate (YC S25, founded 2024, San Francisco) builds managed reinforcement-learning sandbox environments, simulated …
Matrices · Matrices builds reinforcement-learning training environments for frontier AI labs to train agents that use computers and …
BenchFlow · BenchFlow is an early-stage, YC-backed open-source 'environment lab' building evaluation infrastructure and a community …
Collinear · Collinear AI operates a 'Simulation Lab' (SimLab) that builds sandboxed, stateful RL environments simulating enterprise …
Refresh · Refresh (YC X25) builds simulation engines / RL environments with verifiable rewards for coding and computer use, …
Vmax · Vmax is a San Francisco reinforcement-learning startup (founded 2025 by three RL/robotics PhDs from UCL and UPenn) that …
Andromede · Andromede is an early-stage RL data lab that programmatically generates RL environments, tasks, and verifiers from …
Plato · Plato (plato.so, Plato Technologies, Inc.) builds simulated worlds for training and evaluating browser and computer-use …
AIChamp · AIChamp builds custom reinforcement-learning environments and 'Virtual Gym' simulations for training and evaluating …
Habitat Inc · Habitat Inc is an early-stage commercial vendor (2-10 employees, New York HQ) building reinforcement-learning …
Scale AI · Scale AI is the data-labeling and AI-data incumbent that has extended into RL environments, offering simulated web apps, …
Modal · Modal (Modal Labs) is a New York-based, Python-native serverless cloud purpose-built for AI/ML workloads, providing …
Mercor · Mercor is a venture-backed expert-marketplace and AI-training-data company that organizes a network of ~30,000+ domain …
Surge AI · Surge AI is a bootstrapped, high-revenue human-data and RLHF labeling leader serving frontier AI labs, which has …
Prime Intellect · Prime Intellect operates an open-source RL stack - the Environments Hub (2,500+ community RL environments), the …
Daytona · Daytona provides secure, elastic, programmatic sandboxes ('computers') that AI agents and developers can spin up in …
Deeptune · Deeptune was a New York-based startup building managed reinforcement-learning environments ('training gyms') for …
E2B · E2B provides open-source, secure cloud sandboxes (built on Firecracker microVMs) for running AI-generated code and AI …
Runloop · Runloop sells cloud-hosted, isolated micro-VM 'devboxes' plus blueprints, snapshots and benchmark/eval tooling that give …
General Reasoning · General Reasoning is an AI research lab (operating research hub in London; legal entity General Reasoning, Inc. …
Cua · Cua (trycua, YC X25) is open-source MIT-licensed infrastructure for computer-use agents, providing cloud and self-hosted …
Sepal AI · Sepal AI was a YC-backed (S24) San Francisco data-research company that built high-quality training data, expert-graded …
Good Start Labs · Good Start Labs is a 2025 Every spin-out that builds game-based environments to generate reinforcement-learning data and …
Morph · Morph (Morph Labs) provides snapshot-based VM compute for AI agents via its Infinibranch / Liquid Metal technology, …
Turing · Turing is a large AGI-infrastructure and engineering-services company that supplies frontier AI labs with coding data, …
Mercor · Thiel Fellowship, Harvard, Georgetown
Handshake · Palantir, Meta, Scale AI
Scale · Stanford, Coinbase, Y Combinator
Turing · Stanford, Rover, Yahoo
Surge · Google, Meta, Twitter
Snorkel · Stanford, UW, Meta
micro1 · Stanford, Berkeley
AfterQuery · Wharton, Penn, UBC
Mechanize · Epoch AI
Patronus AI · Meta AI, Google, Amazon
Datacurve · Waterloo
Halluminate · Capital One, Meta, Cornell
Pareto · Stanford, Scale, Turing
Proximal · Prime Intellect, Cursor
Aptura · Lazard, McKinsey, EF
Bespoke Labs · Google DeepMind, UC Berkeley
Collinear · Hugging Face, Salesforce, Stanford
Latch · Berkeley, Asimov, Y Combinator
Quesma · Elastic, Sumo Logic, Dynatrace
ReasonCore · Turing, Aera Technology, Oracle
BenchFlow · Terminal, Red Hat
Cua · Microsoft, Twitter
EdotEnv · G-Research, Etched, ETH Zurich
Emulated · AWS, DeepMind, Prime Intellect
General Reasoning · Meta, Conjecture, Aleph Alpha
Markov · Y Combinator, IIT Madras
Metaphi · Waymo, Nuro, Netflix
pre.dev · Penn M&T, Wharton
Tacit Labs · MosaicML, MSR, EF
Ulam · Oxford
Vmax · UCL
Deeptune · Hebbia
Fleet · Mercor, Anthropic, MSL
Huzzle Labs · Turing
Taste Labs · Exa, Palantir, Mercado Libre
Verita AI · HRT, Jane Street, Mercor
ARIMLABS · Akamai
Chakra Labs · Microsoft, L/S, Snap
dmodel · OpenAI, Google Brain, EleutherAI
HUD · METR, AWS, Hume
Idler · Meta, Microsoft, Cornell
Incalmo · CMU, MSR, RSAC Labs