arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

更好、更快、更强:程序化技能学习最能降低智能体成本

Better, Faster, Stronger: Programmatic Skill Learning Best Reduces Agent Cost

Zixi Huang, Xiheng Wang, Andrew Wang, William Jurayj, Bernal Jiménez Gutiérrez, Daniel Khashabi, Nicholas Andrews

arXiv 2608.11338首次发表:更新:

发表机构

Johns Hopkins University(约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出程序化技能学习的SpeedRunner编码智能体,通过分析轨迹重构技能,在三个具身环境中实现最优学习与成本降低,且对分布偏移和环境随机性能保持鲁棒性。

AI 中文摘要

最近,用技能增强大语言模型(LLM)智能体能力的做法已变得普遍。我们研究通过学习技能让智能体适应新领域的成本效益方法,现有研究多关注性能提升而非成本效益,因此对哪些技能学习策略能节省成本知之甚少。我们认为,在所有不同的技能学习方法中,将技能视为程序的方法可实现最佳成本降低。通过确定性执行动作序列,程序增强型智能体能可靠且低成本地达成目标,否则这些目标需要反复试错,还会在长期任务中面临退化行为风险。智能体可在推理时通过逐步发现这些程序并将其用于未来任务来学习。我们假设,即使没有重放或验证,只要智能体能学会分析过往轨迹,这些轨迹就包含足够信号来指导技能学习。为验证我们的主张,我们提出SpeedRunner(一种编码智能体),它可分析轨迹并重构技能以提升未来任务性能。在三个不同的具身环境中,我们表明SpeedRunner始终在学习和成本降低方面达到前沿水平,同时对分布偏移和环境随机性保持鲁棒性。

英文摘要

Recently, the practice of augmenting LLM agent capability with skills has gained prevalence. We explore the cost effective adaptation of agents to novel domains by means of learning skills. Existing works focus on performance gain over cost effectiveness. As a result, little is known about what skill learning strategies save cost. We argue that among all the different skill learning methods, those that view skills as programs can achieve the best cost reduction. By executing sequences of actions deterministically, a program-augmented agent can reliably and cheaply achieve goals that would otherwise require trial and error and risk degenerate behavior over long horizons. An agent can learn at inference time by incrementally discovering these programs and equipping them for future tasks. We hypothesize that past trajectories contain enough signal to guide skill learning, even without replay or validation, provided the agent can learn to analyze them. To test our claims, we propose SpeedRunner, a coding agent that analyzes trajectories and refactors skills for better performance on future tasks. Across three different embodied environments, we show that SpeedRunner consistently achieves the frontier in learning and cost reduction while remaining robust against distribution shifts and environmental randomness.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑