AnyAct:自我进化智能体的通用动作
AnyAct: Universal Action for Self-Evolving Agents
浏览论文内容
中文总结 AI 辅助
AnyAct提出通用动作层,通过分层检索与可靠性进化构建自我进化动作空间,统一异构反馈,在LiveMCPBench和OSMCP上以更少步骤实现最优性能。
中文摘要 AI 辅助
随着大语言模型(LLM)的进步,AI智能体越来越多地被部署在开放世界环境中,以处理复杂的序列任务(例如文档处理、跨应用协作),其高度依赖于从GUI操作到语义API的各种动作。然而,三个核心挑战依然存在:大规模工具生态系统超出LLM上下文窗口的“规模困境”、因更新或中断导致的工具质量的“非平稳性”,以及反馈格式(像素、文本、结构化数据)的“异构性”造成的信息孤岛。为解决这些问题,我们提出了AnyAct,一个通用动作层,将可用能力统一到一个自我进化的动作空间中,使智能体能够在大规模、动态的工具生态系统中高效可靠地运行。AnyAct的核心设计聚焦于两个目标:通过分层渐进式检索(过滤任务相关动作)和测试时可靠性进化(修剪不可靠动作)来构建该动作空间,并通过异构观察接地模块实现可靠性感知的动作编排,该模块统一了多模态反馈。此外,它定义了一个混合动作空间(原始动作+语义动作),并优化了任务成功率与执行成本之间的平衡。在LiveMCPBench和OSMCP(我们为多粒度动作协作开发的新基准)上的评估展示了最先进的性能。AnyAct在LiveMCPBench上,相对于各种LLM基础模型上的基线方法,带来了显著的性能提升,对于原生能力受限的模型,改进尤为显著。在OSMCP上,它仅用50步就实现了77.27%的总体成功率,这是大多数竞争对手所需步数的一半。
英文摘要
As large language models (LLMs) advance, AI agents are increasingly deployed in open-world environments to tackle complex sequential tasks (e.g., document processing, cross-application collaboration), relying heavily on actions ranging from GUI operations to semantic APIs. However, three core challenges persist: the "scale dilemma" of massive tool ecosystems exceeding LLM context windows, the "non-stationarity" of tool quality due to updates or outages, and the "heterogeneity" of feedback formats (pixels, text, structured data) creating information silos. To address these, we propose AnyAct, a universal action layer that unifies available capabilities into a self-evolving action space, enabling agents to operate efficiently and reliably in large-scale, dynamic tool ecosystems. AnyAct's core design focuses on two objectives: constructing this action space via hierarchical progressive retrieval (filtering task-relevant actions) and test-time reliability evolution (pruning unreliable actions), and enabling reliability-aware action orchestration through a heterogeneous observation grounding module that unifies multi-modal feedback. Additionally, it defines a hybrid action space (primitive + semantic actions) and optimizes for a balance between task success rate and execution cost. Evaluations on LiveMCPBench and OSMCP (a new benchmark we developed for multi-granularity action collaboration) demonstrate state-of-the-art performance. AnyAct delivers substantial performance gains over baseline methods across various LLM base models on LiveMCPBench and improvements are particularly notable for models with constrained native capabilities. On OSMCP, it achieves 77.27% overall success with only 50 steps, which is half the steps required by most competitors.
发表机构
- The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。