arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

驾驭演化即学习:自我改进型个人智能体的逼近、泛化与优化极限

Harness Evolution as Learning: Approximation, Generalization, and Optimization Limits of Self-Improving Personal Agents

Zeyu Gan, Zixuan Gong, Yong Liu

arXiv 2609.36892首次发表:更新:

发表机构

Gaoling School of Artificial Intelligence; Renmin University of China(高瓴人工智能学院; 中国人民大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究将个人智能体的驾驭演化形式化为学习问题,通过经验基准和理论分析揭示其逼近、泛化与优化极限,为设计提供统一视角。

AI 中文摘要

随着大型语言模型(LLM)能力的持续进步,越来越多的关注转向如何将其能力转化为有用的行为。个人智能体将这一问题带入日常场景,在此类场景中,模型需要服务个体用户并持续适应其偏好。在底层模型保持不变的情况下,这种适应依赖于驾驭工程(harness engineering):设计和演化管理上下文、记忆、工具和执行的周边层。尽管进展迅速,但支配有效驾驭演化的因素仍未得到充分理解。为缩小这一差距,我们通过互补的经验与理论分析,研究了关于驾驭架构、驾驭规模以及自我演化算法的三个核心问题。在经验层面,我们引入了一个偏好导向的基准,并系统地表征了与这三个维度相关的个人智能体的能力与局限。在理论层面,我们将驾驭演化形式化为一个学习问题,并通过逼近误差、泛化误差和优化误差来解释这些现象。对可达策略、有限交互证据下的容量以及有偏更新动态的分析,为所观察到的现象提供了理论解释。综合来看,这些结果为通过驾驭演化实现个性化的极限提供了一个统一视角,并为未来的驾驭设计提供参考。

英文摘要

As the capabilities of large language models (LLMs) continue to advance, increasing attention is turning to how to translate their abilities into useful behavior. Personal agents bring this question into everyday settings, where models are expected to serve individual users and continually adapt to their preferences. With the underlying model held fixed, such adaptation relies on harness engineering: designing and evolving the surrounding layer that manages context, memory, tools, and execution. Despite rapid progress, the factors governing effective harness evolution remain insufficiently understood. To narrow this gap, we investigate three central questions concerning harness architecture, harness scale, and self-evolution algorithms through complementary empirical and theoretical analyses. Empirically, we introduce a preference-oriented benchmark and systematically characterize the capabilities and limitations of personal agents associated with these three dimensions. Theoretically, we formulate harness evolution as a learning problem and explain these phenomena through approximation, generalization, and optimization errors. Analyses of reachable policies, capacity under finite interaction evidence, and biased update dynamics provide theoretical accounts of the observed phenomena. Together, these results offer a unified perspective on the limits of personalization through harness evolution and inform future harness design.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑