arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

抽象阶梯的上下行:语言代理的基于代码的技能

Up and Down the Abstraction Ladder: Code-Based Skills for Language Agents

Bartłomiej Cupiał, Jens Tuyls, Maciej Wołczyk, Davide Paglieri, Martin Klissarov, Benjamin Eysenbach, Piotr Miłoś, Karthik R. Narasimhan

arXiv 2609.31076首次发表:更新:

发表机构

University of Warsaw; Princeton University; IDEAS NCBR; University College London; McGill University; AKCES NCBR; Mila; Mistral AI; Institute of Mathematics, Polish Academy of Sciences(华沙大学; 普林斯顿大学; IDEAS NCBR; 伦敦大学学院; 麦吉尔大学; AKCES NCBR; 米拉研究所; Mistral AI; 波兰科学院数学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过CodeHack技能库在NetHack环境中系统评估代码抽象对语言代理的影响,发现技能可提升性能、降低推理成本并加速学习,同时保留原始动作以维持灵活性。

AI 中文摘要

语言代理在需要长序列低层动作的环境中难以行动和学习。基于代码的抽象可以通过让代理调用可重用技能,而非反复选择单个动作,来提高其生产力。代码处理重复出现的局部决策,而语言模型决定使用哪些技能以及如何组合它们。然而,抽象是有漏洞的,超出技能能力范围的情况可能需要回归到原始动作。受这种生产力与灵活性之间权衡的驱动,我们系统地研究了基于代码的动作抽象如何影响语言代理的性能、推理成本和学习。我们在NetHack(一个具有挑战性的、长期视野的游戏环境)中,使用CodeHack(我们的基于代码的技能库,附带自然语言描述)对此进行研究。我们使用该库比较了仅限于原始动作的代理与仅使用语义技能或结合原始动作使用技能的代理。我们在三种设置中评估这些代理:零样本提示、监督微调和强化学习。在NetHack上的广泛零样本评估中,我们发现与原始动作相比,技能几乎使游戏进度提高了三倍,同时将每集的推理成本降低了86%。将技能与原始动作结合保留了大部分收益,同时保留了返回低层动作的路径。最后,在强化学习中,我们发现基于技能的代理比基于原始动作的代理学习速度快得多,在相同的训练预算下,地牢关卡的平均增益提高了7.2倍。这些结果表明,提供的技能库可以提高性能、效率和学习,而保留原始动作则在技能库不足时提供灵活性。我们发布了CodeHack以及训练和评估代码。

英文摘要

Language agents struggle to act and learn in environments that require long sequences of low-level actions. Code-based abstractions can make these agents more productive by letting them invoke reusable skills instead of repeatedly selecting individual actions. The code handles recurring local decisions, while the language model decides which skills to use and how to combine them. Yet abstractions are leaky, and situations beyond a skill's capabilities may require a return to primitive actions. Motivated by this tradeoff between productivity and flexibility, we systematically study how code-based action abstraction affects the performance, inference cost, and learning of language agents. We study this in NetHack, a challenging, long-horizon game environment, using CodeHack, our library of code-based skills with natural-language descriptions. We use this library to compare agents restricted to primitives with those using semantic skills alone or in combination with primitives. We evaluate these agents in three settings: zero-shot prompting, supervised fine-tuning, and reinforcement learning. Across a broad zero-shot evaluation on NetHack, we find that compared with primitives, skills nearly triple game progression, while reducing inference cost per episode by 86%. Combining skills with primitives retains much of this benefit while preserving a path back down to low-level actions. Finally, in RL, we find that skill-based agents learn significantly faster than agents acting on primitives, achieving a 7.2x larger average gain in dungeon level over the same training budget. These results show that a supplied skill library can improve performance, efficiency, and learning, while retaining primitives provides flexibility when the library is insufficient. We release CodeHack together with training and evaluation code.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑