Coding: RL environments and AI benchmarks
35 environments · 26 benchmarks · snapshot 2026-09-03
Scroll to zoom · drag to pan · ○ environment · ◇ benchmark · ● domain
RL environments 35
- AfterQuery
- Mercor
- Prime Intellect
- Mechanize
- Deeptune
- Turing
- Scale AI
- Modal
- Scale
- Morph
- BenchFlow
- General Reasoning
- Vmax
- Datacurve
- Proximal
- Huzzle Labs
- Collinear
- Refresh
- Runloop
- Habitat Inc
- Daytona
- E2B
- Snorkel
- Patronus AI
- ReasonCore
- Abundant
- Emulated
- Metaphi
- pre.dev
- Idler
- Preference Model
- Akhara
- Exabite
- Originator
- Vetto AI
AI benchmarks 26
- SWE-bench Pro
- Terminal-Bench / Terminal-Bench 2.0
- Senior SWE-Bench
- FinanceQA
- IDE-Bench
- APEX
- APEX-Agents
- LiveCodeBench-Plus
- SWE-bench Docker
- SWE-bench Extra
- SWE-bench JavaScript
- SWE-bench Verified
- LiveCodeBench
- Terminal-Bench 2.0
- Terminal-Bench 2.1
- SWE-bench Verified Mini
- FinanceBench
- AA LiveCodeBench
- AA Terminal-Bench 2.1
- HumanEval
- Codeforces
- LiveCodeBench v6
- LiveCodeBench v5
- LiveCodeBench Pass@1-COT
- LiveCodeBench Pro
- Multi-SWE Bench
Next: pick another domain. Tags come from the RL Engineering daily scrape, canonicalized nightly.