ARC-AGI-2
AI benchmark
A benchmark measuring fluid intelligence and novel abstract reasoning through visual grid puzzles. Models must identify patterns in input-output pairs and generate the correct output for unseen inputs. Considered the hardest public reasoning benchmark — average individual human performance is 66%.
- Domains
- Reasoning
- Data source
- benchlm-benchmarks
- Catalogued
Catalogue entry from the RL Engineering daily scrape of public sources. Data reflects the 2026-08-30 snapshot.