RL engineering intelligence, updated daily.

A managed leaderboard of RL environments, benchmarks, and model scores — scraped, normalized, and synthesized by agents from public research sources.

Snapshot 2026-08-17 benchlm-benchmarks benchlm-models benchlm-tools pavlovslist rl-list

Daily briefing

The leaderboard now tracks 535 benchmarks across 98 environments and companies, with 291 models and 9 tools. Notable new benchmarks include GBA Eval for coding Game Boy Advance emulators, VADER for vulnerability assessment, and FinanceQA for financial analysis in LLMs. Key companies featured are Mechanize, AfterQuery, and Bespoke Labs, with tools like Bespoke Curator for synthetic data curation and Evalchemy for evaluation workflows gaining attention. The OpenThoughts dataset collaboration with DataComp highlights ongoing advances in open reasoning benchmarks.

Live GCS test

Loading from /api/site

535
benchmarks
98
companies
98
environments
291
models
9
tools

Model leaderboard

Top BenchLM model scores from the latest scrape. Scores parsed from public leaderboard data.

RankModelScoreCreator
1 Claude Mythos 583.21Anthropic
2 Claude Opus 583.07Anthropic
3 Claude Fable 582.96Anthropic
4 GPT-5.6 Sol82.00OpenAI
5 Kimi K380.50Moonshot AI
6 Qwen3.8 Max79.91Alibaba
7 Claude Opus 4.876.58Anthropic
8 Gemini 3.6 Flash75.54Google
9 Grok 4.575.42xAI
10 GPT-5.573.37OpenAI
11 GPT-5.473.36OpenAI
12 Claude Opus 4.771.87Anthropic
13 Qwen3.7 Max71.60Alibaba
14 MiniMax M368.74MiniMax
15 Hy368.02Tencent
16 Claude Opus 4.667.98Anthropic
17 Gemini 3 Pro67.37Google
18 Inkling67.04Thinking Machines Lab
19 MiMo-V2-Pro66.94Xiaomi
20 GLM-5.166.91Z.AI
21 Qwen3.7 Plus66.02Alibaba
22 GPT-5.3 Codex65.93OpenAI
23 GLM-565.61Z.AI
24 Gemini 3.5 Flash-Lite65.29Google
25 Qwen3.6 Plus64.79Alibaba
26 Claude Sonnet 564.78Anthropic
27 Gemini 3.5 Flash64.67Google
28 Claude Sonnet 4.664.53Anthropic
29 Grok 4.363.95xAI
30 Claude Opus 4.563.63Anthropic
31 Grok 4.663.41xAI
32 GLM-5.263.34Z.AI
33 MiniMax M2.763.06MiniMax
34 MiMo-V2-Omni62.36Xiaomi
35 Muse Spark 1.261.67Meta
36 Gemini 3.7 Flash61.36Google
37 DeepSeek V4 Pro 081361.16DeepSeek
38 Agents-A160.79InternScience
39 GLM-4.760.66Z.AI
40 Grok 4.160.43xAI
41 Kimi K2.660.21Moonshot AI
42 Gemma 4 31B60.15Google
43 Qwen 3.6 Max (preview)59.99Alibaba
44 Qwen3.5-27B59.82Alibaba
45 Gemini 3 Flash59.74Google
46 Qwen3.5-122B-A10B59.64Alibaba
47 Grok 459.61xAI
48 GPT-5.3 Instant59.36OpenAI
49 MiMo-V2.559.16Xiaomi
50 GPT-5 (high)59.08OpenAI

Companies

RL environment vendors and research orgs tracked across sources.

Mechanize

Mechanize is a small, elite San Francisco vendor (founded April 2025 by ex-Epoch AI researchers Matthew Barnett, Tamay Besiroglu, and Ege Erdil) that builds a …

rl-list San Francisco, USA

AfterQuery

AfterQuery is a San Francisco applied-research lab and data platform (YC W25) that supplies frontier AI labs with expert-generated human data (SFT, RL rubrics), …

rl-list San Francisco, USA

Bespoke Labs

Bespoke Labs is an applied AI research lab (Mountain View, CA, founded 2024) focused on data curation and RL-environment curation for training and evaluating …

rl-list Mountain View, California, USA

Huzzle Labs

Huzzle Labs is the AI division of London-based talent platform Huzzle (founded ~2020 by Ingmar Klein, Parham Rakhshanfar, and Amit Choudhary). It positions …

rl-list London, United Kingdom

Fleet AI

Fleet AI builds high-fidelity reinforcement-learning training environments ('gyms') that replicate enterprise software such as Salesforce and Excel, plus …

rl-list New York, NY, USA

Datacurve

Datacurve is a YC W24 commercial data vendor that supplies expert-curated frontier coding data, RLHF traces, and repository-wide reinforcement learning …

rl-list San Francisco, USA

Proximal

Proximal is a San Francisco-based (with a Bangalore presence) research lab for coding data, building high-fidelity, long-horizon reinforcement learning …

rl-list San Francisco, CA, USA

Gray Swan AI

Gray Swan AI is a Pittsburgh-based AI security company spun out of Carnegie Mellon, offering adversarial red-teaming and runtime protection for AI models and …

rl-list Pittsburgh, Pennsylvania, USA

Veris AI

Veris AI sells a high-fidelity simulation platform plus a production runtime that let enterprises train, evaluate, and govern AI agents against mocked …

rl-list San Francisco, CA, USA

Chakra Labs

Chakra Labs runs Dojo, an open/collaborative reinforcement-learning environment hub for computer-use agents, offering deterministic, frame-accurate clones of …

rl-list Brooklyn, New York, USA

Andon Labs

Andon Labs is a Y Combinator-backed (W24) startup, formerly Vectorview, building benchmarks and evaluations for AI agents' long-horizon coherence and safety …

rl-list San Francisco, USA

HUD

HUD (YC W25, formerly hud.so) is a platform for building reinforcement-learning environments and evaluations for computer-use and browser agents. It lets teams …

rl-list San Francisco, USA

Vals AI

Vals AI is an independent, third-party benchmarking and evaluation platform that scores LLMs and AI applications (copilots, RAG, agents) on rigorous, …

rl-list San Francisco, USA

Halluminate

Halluminate (YC S25, founded 2024, San Francisco) builds managed reinforcement-learning sandbox environments, simulated applications, and human/annotation data …

rl-list San Francisco, CA, USA

Matrices

Matrices builds reinforcement-learning training environments for frontier AI labs to train agents that use computers and browsers like humans, described as a …

rl-list San Francisco, California, USA

BenchFlow

BenchFlow is an early-stage, YC-backed open-source 'environment lab' building evaluation infrastructure and a community Benchmark Hub for AI agents, with …

rl-list New Castle, DE, USA (incorporation); Bay Area / San Francisco operating presence

Collinear

Collinear AI operates a 'Simulation Lab' (SimLab) that builds sandboxed, stateful RL environments simulating enterprise users, tools (Jira, ServiceNow, Shopify, …

rl-list Mountain View / Sunnyvale, California, USA

Refresh

Refresh (YC X25) builds simulation engines / RL environments with verifiable rewards for coding and computer use, partnering with frontier labs and enterprises …

rl-list San Francisco, CA, USA

Vmax

Vmax is a San Francisco reinforcement-learning startup (founded 2025 by three RL/robotics PhDs from UCL and UPenn) that automates the conversion of proprietary …

rl-list San Francisco, USA

Andromede

Andromede is an early-stage RL data lab that programmatically generates RL environments, tasks, and verifiers from real-world data for post-training and …

rl-list Lausanne, Switzerland

Plato

Plato (plato.so, Plato Technologies, Inc.) builds simulated worlds for training and evaluating browser and computer-use agents, recreating real …

rl-list San Francisco, CA, USA

AIChamp

AIChamp builds custom reinforcement-learning environments and 'Virtual Gym' simulations for training and evaluating tool-using AI agents on long-horizon, …

rl-list San Francisco, USA (CEO-based; reported, not officially confirmed; Tracxn alternatively lists Bali, Indonesia)

Habitat Inc

Habitat Inc is an early-stage commercial vendor (2-10 employees, New York HQ) building reinforcement-learning environments for white-collar / work automation, …

rl-list New York, NY, USA

Scale AI

Scale AI is the data-labeling and AI-data incumbent that has extended into RL environments, offering simulated web apps, macOS/Windows-like desktop VMs, and …

rl-list San Francisco, California, USA

Modal

Modal (Modal Labs) is a New York-based, Python-native serverless cloud purpose-built for AI/ML workloads, providing on-demand GPU/CPU compute, fast-booting …

rl-list New York, NY, USA

Mercor

Mercor is a venture-backed expert-marketplace and AI-training-data company that organizes a network of ~30,000+ domain experts (doctors, lawyers, bankers, …

rl-list San Francisco, CA, USA (181 Fremont)

Surge AI

Surge AI is a bootstrapped, high-revenue human-data and RLHF labeling leader serving frontier AI labs, which has expanded into agentic RL environments via its …

rl-list San Francisco, California, USA

Prime Intellect

Prime Intellect operates an open-source RL stack - the Environments Hub (2,500+ community RL environments), the Verifiers library and prime-rl training …

rl-list San Francisco, USA

Daytona

Daytona provides secure, elastic, programmatic sandboxes ('computers') that AI agents and developers can spin up in under ~90ms to run untrusted AI-generated …

rl-list New York, NY, United States

Deeptune

Deeptune was a New York-based startup building managed reinforcement-learning environments ('training gyms') for computer-use and code, where AI agents practice …

rl-list New York, NY, USA

E2B

E2B provides open-source, secure cloud sandboxes (built on Firecracker microVMs) for running AI-generated code and AI agents, offered as a hosted API with …

rl-list San Francisco, USA

Runloop

Runloop sells cloud-hosted, isolated micro-VM 'devboxes' plus blueprints, snapshots and benchmark/eval tooling that give AI coding agents a secure execution …

rl-list San Francisco, CA, USA

General Reasoning

General Reasoning is an AI research lab (operating research hub in London; legal entity General Reasoning, Inc. registered in the US) building open RL …

rl-list London, United Kingdom (Shoreditch), operating research hub; legal entity General Reasoning, Inc. registered in San Francisco/US per SEC Form D

Cua

Cua (trycua, YC X25) is open-source MIT-licensed infrastructure for computer-use agents, providing cloud and self-hosted sandboxes across macOS, Windows, Linux, …

rl-list San Francisco, CA, USA

Sepal AI

Sepal AI was a YC-backed (S24) San Francisco data-research company that built high-quality training data, expert-graded evaluation benchmarks, and …

rl-list San Francisco, USA

Good Start Labs

Good Start Labs is a 2025 Every spin-out that builds game-based environments to generate reinforcement-learning data and evaluate frontier models, using both …

rl-list Brooklyn, NY, USA

Morph

Morph (Morph Labs) provides snapshot-based VM compute for AI agents via its Infinibranch / Liquid Metal technology, which can snapshot, branch, and restore …

rl-list San Francisco, USA

Turing

Turing is a large AGI-infrastructure and engineering-services company that supplies frontier AI labs with coding data, human expertise, and RL/evaluation data …

rl-list Palo Alto, USA

Mercor

Thiel Fellowship, Harvard, Georgetown

pavlovslist SFSanta Clara

Handshake

Palantir, Meta, Scale AI

pavlovslist SFNYC, Bangalore, Berlin

Scale

Stanford, Coinbase, Y Combinator

pavlovslist SFNYC, DC, London

Turing

Stanford, Rover, Yahoo

pavlovslist SFPalo Alto, Gurugram

Surge

Google, Meta, Twitter

pavlovslist SFNYC, Seattle

Snorkel

Stanford, UW, Meta

pavlovslist SFRedwood City, NYC

micro1

Stanford, Berkeley

pavlovslist SF

AfterQuery

Wharton, Penn, UBC

pavlovslist SFNYC, Seattle

Pareto

Stanford, Scale, Turing

pavlovslist SF

Proximal

Prime Intellect, Cursor

pavlovslist SFBangalore

Aptura

Lazard, McKinsey, EF

pavlovslist LondonSF

Bespoke Labs

Google DeepMind, UC Berkeley

pavlovslist Mountain ViewMenlo Park, Bangalore, SF

Collinear

Hugging Face, Salesforce, Stanford

pavlovslist Mountain ViewSunnyvale

Latch

Berkeley, Asimov, Y Combinator

pavlovslist SF

Quesma

Elastic, Sumo Logic, Dynatrace

pavlovslist Warsaw

ReasonCore

Turing, Aera Technology, Oracle

pavlovslist SF

Cua

Microsoft, Twitter

pavlovslist SF

EdotEnv

G-Research, Etched, ETH Zurich

pavlovslist SF

Emulated

AWS, DeepMind, Prime Intellect

pavlovslist SF

Markov

Y Combinator, IIT Madras

pavlovslist SF

Metaphi

Waymo, Nuro, Netflix

pavlovslist SFNYC

pre.dev

Penn M&T, Wharton

pavlovslist Delaware

Ulam

Oxford

pavlovslist WarsawLondon

Vmax

UCL

pavlovslist SFNYC

Fleet

Mercor, Anthropic, MSL

pavlovslist SFNYC

Taste Labs

Exa, Palantir, Mercado Libre

pavlovslist NYC

Verita AI

HRT, Jane Street, Mercor

pavlovslist SFVancouver

dmodel

OpenAI, Google Brain, EleutherAI

pavlovslist SF

HUD

METR, AWS, Hume

pavlovslist SFSingapore

Idler

Meta, Microsoft, Cornell

pavlovslist SF

Incalmo

CMU, MSR, RSAC Labs

pavlovslist San Mateo

Benchmarks

Public benchmarks and eval suites referenced by tracked companies.

APEX-Agents

Mercor · Coding, Enterprise Workflows, Long-Horizon

rl-list

APEX-SWE

Mercor · Coding, Enterprise Workflows, Long-Horizon

rl-list

APEX

Mercor · Multi-Domain

pavlovslist

Realm

micro1 · Multi-Domain

pavlovslist

AppBench

AfterQuery · Multi-Domain

pavlovslist

DeepSWE

Datacurve · Code

pavlovslist

WebBench

Halluminate · Finance, Enterprise

pavlovslist

Leap

Pareto · Multi-Domain, Tool Use

pavlovslist

Environments

RL environments and tooling surfaces discovered across the ecosystem.

Mechanize

Mechanize · Mechanize is a small, elite San Francisco vendor (founded April 2025 by ex-Epoch AI researchers Matthew Barnett, Tamay …

rl-list

AfterQuery

AfterQuery · AfterQuery is a San Francisco applied-research lab and data platform (YC W25) that supplies frontier AI labs with …

rl-list

Bespoke Labs

Bespoke Labs · Bespoke Labs is an applied AI research lab (Mountain View, CA, founded 2024) focused on data curation and RL-environment …

rl-list

Huzzle Labs

Huzzle Labs · Huzzle Labs is the AI division of London-based talent platform Huzzle (founded ~2020 by Ingmar Klein, Parham …

rl-list

Fleet AI

Fleet AI · Fleet AI builds high-fidelity reinforcement-learning training environments ('gyms') that replicate enterprise software …

rl-list

Datacurve

Datacurve · Datacurve is a YC W24 commercial data vendor that supplies expert-curated frontier coding data, RLHF traces, and …

rl-list

Proximal

Proximal · Proximal is a San Francisco-based (with a Bangalore presence) research lab for coding data, building high-fidelity, …

rl-list

Gray Swan AI

Gray Swan AI · Gray Swan AI is a Pittsburgh-based AI security company spun out of Carnegie Mellon, offering adversarial red-teaming and …

rl-list

Veris AI

Veris AI · Veris AI sells a high-fidelity simulation platform plus a production runtime that let enterprises train, evaluate, and …

rl-list

Chakra Labs

Chakra Labs · Chakra Labs runs Dojo, an open/collaborative reinforcement-learning environment hub for computer-use agents, offering …

rl-list

Andon Labs

Andon Labs · Andon Labs is a Y Combinator-backed (W24) startup, formerly Vectorview, building benchmarks and evaluations for AI …

rl-list

HUD

HUD · HUD (YC W25, formerly hud.so) is a platform for building reinforcement-learning environments and evaluations for …

rl-list

Vals AI

Vals AI · Vals AI is an independent, third-party benchmarking and evaluation platform that scores LLMs and AI applications …

rl-list

Halluminate

Halluminate · Halluminate (YC S25, founded 2024, San Francisco) builds managed reinforcement-learning sandbox environments, simulated …

rl-list

Matrices

Matrices · Matrices builds reinforcement-learning training environments for frontier AI labs to train agents that use computers and …

rl-list

BenchFlow

BenchFlow · BenchFlow is an early-stage, YC-backed open-source 'environment lab' building evaluation infrastructure and a community …

rl-list

Collinear

Collinear · Collinear AI operates a 'Simulation Lab' (SimLab) that builds sandboxed, stateful RL environments simulating enterprise …

rl-list

Refresh

Refresh · Refresh (YC X25) builds simulation engines / RL environments with verifiable rewards for coding and computer use, …

rl-list

Vmax

Vmax · Vmax is a San Francisco reinforcement-learning startup (founded 2025 by three RL/robotics PhDs from UCL and UPenn) that …

rl-list

Andromede

Andromede · Andromede is an early-stage RL data lab that programmatically generates RL environments, tasks, and verifiers from …

rl-list

Plato

Plato · Plato (plato.so, Plato Technologies, Inc.) builds simulated worlds for training and evaluating browser and computer-use …

rl-list

AIChamp

AIChamp · AIChamp builds custom reinforcement-learning environments and 'Virtual Gym' simulations for training and evaluating …

rl-list

Habitat Inc

Habitat Inc · Habitat Inc is an early-stage commercial vendor (2-10 employees, New York HQ) building reinforcement-learning …

rl-list

Scale AI

Scale AI · Scale AI is the data-labeling and AI-data incumbent that has extended into RL environments, offering simulated web apps, …

rl-list

Modal

Modal · Modal (Modal Labs) is a New York-based, Python-native serverless cloud purpose-built for AI/ML workloads, providing …

rl-list

Mercor

Mercor · Mercor is a venture-backed expert-marketplace and AI-training-data company that organizes a network of ~30,000+ domain …

rl-list

Surge AI

Surge AI · Surge AI is a bootstrapped, high-revenue human-data and RLHF labeling leader serving frontier AI labs, which has …

rl-list

Prime Intellect

Prime Intellect · Prime Intellect operates an open-source RL stack - the Environments Hub (2,500+ community RL environments), the …

rl-list

Daytona

Daytona · Daytona provides secure, elastic, programmatic sandboxes ('computers') that AI agents and developers can spin up in …

rl-list

Deeptune

Deeptune · Deeptune was a New York-based startup building managed reinforcement-learning environments ('training gyms') for …

rl-list

E2B

E2B · E2B provides open-source, secure cloud sandboxes (built on Firecracker microVMs) for running AI-generated code and AI …

rl-list

Runloop

Runloop · Runloop sells cloud-hosted, isolated micro-VM 'devboxes' plus blueprints, snapshots and benchmark/eval tooling that give …

rl-list

General Reasoning

General Reasoning · General Reasoning is an AI research lab (operating research hub in London; legal entity General Reasoning, Inc. …

rl-list

Cua

Cua · Cua (trycua, YC X25) is open-source MIT-licensed infrastructure for computer-use agents, providing cloud and self-hosted …

rl-list

Sepal AI

Sepal AI · Sepal AI was a YC-backed (S24) San Francisco data-research company that built high-quality training data, expert-graded …

rl-list

Good Start Labs

Good Start Labs · Good Start Labs is a 2025 Every spin-out that builds game-based environments to generate reinforcement-learning data and …

rl-list

Morph

Morph · Morph (Morph Labs) provides snapshot-based VM compute for AI agents via its Infinibranch / Liquid Metal technology, …

rl-list

Turing

Turing · Turing is a large AGI-infrastructure and engineering-services company that supplies frontier AI labs with coding data, …

rl-list

Mercor

Mercor · Thiel Fellowship, Harvard, Georgetown

pavlovslist

Handshake

Handshake · Palantir, Meta, Scale AI

pavlovslist

Scale

Scale · Stanford, Coinbase, Y Combinator

pavlovslist

Turing

Turing · Stanford, Rover, Yahoo

pavlovslist

Surge

Surge · Google, Meta, Twitter

pavlovslist

Snorkel

Snorkel · Stanford, UW, Meta

pavlovslist

micro1

micro1 · Stanford, Berkeley

pavlovslist

AfterQuery

AfterQuery · Wharton, Penn, UBC

pavlovslist

Patronus AI

Patronus AI · Meta AI, Google, Amazon

pavlovslist

Halluminate

Halluminate · Capital One, Meta, Cornell

pavlovslist

Pareto

Pareto · Stanford, Scale, Turing

pavlovslist

Proximal

Proximal · Prime Intellect, Cursor

pavlovslist

Aptura

Aptura · Lazard, McKinsey, EF

pavlovslist

Bespoke Labs

Bespoke Labs · Google DeepMind, UC Berkeley

pavlovslist

Collinear

Collinear · Hugging Face, Salesforce, Stanford

pavlovslist

Latch

Latch · Berkeley, Asimov, Y Combinator

pavlovslist

Quesma

Quesma · Elastic, Sumo Logic, Dynatrace

pavlovslist

ReasonCore

ReasonCore · Turing, Aera Technology, Oracle

pavlovslist

BenchFlow

BenchFlow · Terminal, Red Hat

pavlovslist

Cua

Cua · Microsoft, Twitter

pavlovslist

EdotEnv

EdotEnv · G-Research, Etched, ETH Zurich

pavlovslist

Emulated

Emulated · AWS, DeepMind, Prime Intellect

pavlovslist

Markov

Markov · Y Combinator, IIT Madras

pavlovslist

Metaphi

Metaphi · Waymo, Nuro, Netflix

pavlovslist

pre.dev

pre.dev · Penn M&T, Wharton

pavlovslist

Tacit Labs

Tacit Labs · MosaicML, MSR, EF

pavlovslist

Ulam

Ulam · Oxford

pavlovslist

Vmax

Vmax · UCL

pavlovslist

Deeptune

Deeptune · Hebbia

pavlovslist

Fleet

Fleet · Mercor, Anthropic, MSL

pavlovslist

Taste Labs

Taste Labs · Exa, Palantir, Mercado Libre

pavlovslist

Verita AI

Verita AI · HRT, Jane Street, Mercor

pavlovslist

ARIMLABS

ARIMLABS · Akamai

pavlovslist

Chakra Labs

Chakra Labs · Microsoft, L/S, Snap

pavlovslist

dmodel

dmodel · OpenAI, Google Brain, EleutherAI

pavlovslist

HUD

HUD · METR, AWS, Hume

pavlovslist

Idler

Idler · Meta, Microsoft, Cornell

pavlovslist

Incalmo

Incalmo · CMU, MSR, RSAC Labs

pavlovslist