Long-Horizon Terminal-Bench
AI benchmark from Paper Arxiv
Long-Horizon Terminal-Bench evaluates whether agents can sustain progress on 46 reproducible terminal tasks spanning nine categories, including software engineering, scientific computing, multimodal analysis, research reproduction, systems, professional workflows, games, and logic puzzles. Tasks run in containers under a 90-minute budget and use hidden replay-based verifiers with continuous partial credit, stressing long-horizon planning, context management, iterative debugging, and recovery across hundreds of dependent actions.
- Publisher
- Paper Arxiv
- Domains
- Agentic
- Data source
- benchmarklist
- Catalogued
Catalogue entry from the RL Engineering daily scrape of public sources. Data reflects the 2026-08-30 snapshot.