arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

KnowAct-GUIClaw:深度理解、完美执行,具有自我进化记忆和技能的个人GUI助手

KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill

Yunxin Li, Jinchao Li, Shibo Su, Zhenran Xu, Chenrui Zhao, Tongshu Bian, Xiaoman Liang, Meishan Zhang, Baotian Hu, Min Zhang

arXiv 2607.12625首次发表:更新:

发表机构

Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对OpenClaw的不足,提出深度理解、完美执行范式,介绍KnowAct-GUIClaw框架,通过多系统实验验证其在效率、准确性和跨平台适应性方面的优势,且知识记忆和执行技能可跨基础模型迁移。

AI 中文摘要

OpenClaw作为复杂任务自动化的领先代理框架,存在跨平台GUI交互支持不足和自我进化机制不完善的问题,限制了其在不同设备生态系统中的应用及性能提升。本文提出个人助手的深度理解、完美执行范式,介绍了KnowAct-GUIClaw这一新颖框架,它能解决OpenClaw的GUI操作缺陷,突破跨平台和递归自我改进约束。通过在多系统上的广泛实验表明,KnowAct-GUIClaw具有卓越的效率、准确性和跨平台适应性,知识记忆和执行技能可跨多种基础模型迁移。

英文摘要

OpenClaw has emerged as a leading agent framework for complex task automation, yet it faces insufficient cross-platform GUI interaction support and a well-built self-evolution mechanism. These flaws limit its adaptation to diverse device ecosystems and prevent performance improvements through continuous learning from execution experience. To resolve these issues, we propose the Know Deeply, Act Perfectly paradigm for personal assistants, which holds that accumulated user interaction and task-running experience directly improve execution accuracy and efficiency, unifying cognitive comprehension and operational execution. Based on this paradigm, we introduce KnowAct-GUIClaw, a novel Know-Route-Act-Reflect framework designed to address OpenClaw's GUI manipulation deficits and break through its cross-platform and recursive self-improvement constraints. First, the host agent leverages accumulated interaction experience and task-relevant knowledge for long-horizon task decomposition and allocation (Know). Second, a pluggable GUI subagent with an experience-attributable memory system (Know) and self-evolving skill library (Act), enabling seamless cross-platform migration and fast-path integration. Especially, this framework continuously stores user profiles and feedback to improve the accuracy of task decomposition and tool calls. Extensive experiments across Android, iOS, HarmonyOS and Windows show that KnowAct-GUIClaw achieves superior efficiency, accuracy and cross-platform adaptability. Especially, the GUIClaw with open-source Kimi-2.6 models achieves the best performance (64.1%) on the long-horizon MobileWorld benchmark, beating all agentical frameworks and closed-source agentical models, e.g., Seed-2.0-Pro and GPT-5.5. Additionally, the knowledgeable memory and execution skills supported by our framework are transferable across diverse base models, improving by 8.5% with Kimi-2.6.

Comments29 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑