Long-Horizon Terminal-Bench

AI benchmark from Paper Arxiv

Long-Horizon Terminal-Bench evaluates whether agents can sustain progress on 46 reproducible terminal tasks spanning nine categories, including software engineering, scientific computing, multimodal analysis, research reproduction, systems, professional workflows, games, and logic puzzles. Tasks run in containers under a 90-minute budget and use hidden replay-based verifiers with continuous partial credit, stressing long-horizon planning, context management, iterative debugging, and recovery across hundreds of dependent actions.

Publisher
Paper Arxiv
Domains
Agentic
Data source
benchmarklist
Catalogued
Website
https://benchmarklist.com/benchmarks/long_horizon_terminal_bench/

Catalogue entry from the RL Engineering daily scrape of public sources. Data reflects the 2026-08-30 snapshot.