HumanEval
AI benchmark
A set of 164 handwritten Python function-generation problems. HumanEval is useful as a historical floor check, but BenchLM's current exact-source table is too small to support a broad frontier-coding verdict.
- Domains
- Coding
- Data source
- benchlm-benchmarks
- Catalogued
Catalogue entry from the RL Engineering daily scrape of public sources. Data reflects the 2026-08-30 snapshot.