Multi-SWE Bench
AI benchmark
A multi-language software-engineering benchmark that measures repository-level bug fixing and implementation across more than one programming ecosystem.
- Domains
- Coding
- Data source
- benchlm-benchmarks
- Catalogued
Catalogue entry from the RL Engineering daily scrape of public sources. Data reflects the 2026-08-30 snapshot.