AI
Long-Horizon-Terminal-Bench: 密な報酬ベースの評価による長期間の端末タスクにおけるエージェントの限界のテスト
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
arxiv2026年7月13日
日本語要約
この研究は、密な報酬ベースの評価を使用した長期間の端末タスクにおけるAIエージェントの限界をテストするために設計された新しいベンチマークであるLong-Horizon-Terminal-Benchを紹介します。これは、長期間にわたる複雑で多段階のタスクを実行するエージェントの能力を評価する上での重要なギャップに対処します。
English Summary
This research introduces Long-Horizon-Terminal-Bench, a new benchmark designed to test the limits of AI agents on long-horizon terminal tasks using dense reward-based grading. It addresses a critical gap in evaluating agents' ability to perform complex, multi-step tasks over extended periods.