arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

终端智能体:命令行环境下的AI智能体综述

Terminal Agents: A Survey of AI Agents in Command-Line Environments

Yi Bin, Xiaoyang Yuan, Haoxi Zeng, Wencheng Ye, Wenqi Shao, Chen Qian, Wei Ye, Yujuan Ding, Zheng Wang, Pengpeng Zeng, Jingkuan Song, Heng Tao Shen

arXiv 2608.20485首次发表:更新:

发表机构

Shanghai Innovation Institute; Shanghai Jiao Tong University; The Hong Kong Polytechnic University; Tongji University(上海创新研究院; 上海交通大学; 香港理工大学; 同济大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文将终端智能体定义为行动-观察循环由终端介导的系统,通过七维轮廓关联架构等模块,揭示其行为影响因素与评估不均衡问题,为相关研究提供统一框架。

AI 中文摘要

大型语言模型智能体越来越多地通过终端开展行动,但现有综述将终端介导的行为分散在软件工程、工具使用和计算机使用研究中。本文将终端智能体定义为:其主导的、承载进展的行动-观察循环由终端命令执行、文本反馈及有状态环境交互介导的系统。本文以终端介导的执行为组织视角,划定了工作负载层面的边界,并通过七维终端能力轮廓将系统架构、能力获取与评估关联起来。综合分析表明,实际行为由模型、接口、测试框架(harness)、运行时及环境共同塑造;可执行轨迹将学习锚定在行动后果、验证与恢复中,而现有主流评估侧重最终结果,暴露出过程质量、恢复及治理方面的不均衡。受限固定条件诊断揭示了两点启示:基准系列会呈现不同的过程信号,匹配的系统比较会显示依赖基准的性能及组件归因的局限性。这些发现推动明确报告系统与运行时条件,辅以可复现轨迹与过程层面证据,该框架为跨软件工程及新兴应用领域研究终端介导的智能性提供了统一基础。

英文摘要

Large language model agents increasingly act through terminals, yet existing surveys disperse terminal-mediated behavior across software engineering, tool use, and computer-use research. We regard terminal agents as systems whose dominant progress-bearing action--observation loop is mediated by terminal command execution, textual feedback, and stateful environment interaction. Using terminal-mediated execution as an organizing lens, this survey establishes workload-level boundaries and connects system architecture, competence acquisition, and evaluation through a seven-dimensional terminal competence profile. Our synthesis shows that realized behavior is jointly shaped by the model, interface, harness, runtime, and environment. Executable trajectories ground learning in action consequences, verification, and recovery, whereas prevailing evaluations emphasize final outcomes and expose process quality, recovery, and governance unevenly. Bounded fixed-condition diagnostics illustrate two implications: benchmark families expose different process signals, and matched system comparisons reveal benchmark-dependent performance and limits of component attribution. These findings motivate explicit reporting of system and runtime conditions, supported by replayable traces and process-level evidence. The framework provides a unified basis for studying terminal-mediated agency across software engineering and emerging application domains.

Comments52 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑