arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38885cs.SE

用更少的令牌做更多的事:面向高效编码智能体的分层强化学习

Doing More with Less Tokens: Hierarchical Reinforcement Learning for Efficient Coding Agents

Haobin Li, Liang Jiang, Zhenyu Huang, Mouxing Yang, Xi Peng

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出HERO框架,通过分层强化学习在策略优化中优先任务解决并鼓励高效推理,在SWE基准上实现了解决率与令牌效率的良好平衡。

中文摘要 AI 辅助

近年来,编码智能体已成为现实世界软件工程(SWE)场景中的主导范式,它们通过与开发环境的多轮交互来解决复杂任务。然而,与环境的频繁交互不可避免地会引入大量的令牌开销,导致高昂的使用成本和延迟。尽管近期研究通过推理时的上下文操控和交互限制来探索减少令牌使用,但这些方法侧重于提高令牌效率,却忽视了丢弃任务相关信息所带来的风险,因此难以在解决率与令牌效率之间取得平衡。本文研究了一种不受此限制的更通用范式,即训练具有良好解决性能的令牌高效编码智能体,这是一个高度实用但探索较少的问题。为此,我们揭示了SWE场景中的两个核心观察:i)效率变化:成功解决可以用更少的令牌实现;ii)熵相关性:低效行为与回合级熵相关。受这些观察启发,我们提出了一种新颖的强化学习框架,命名为HERO。具体而言,HERO在策略优化过程中优先考虑任务解决而非令牌效率,并在轨迹和回合两个层面鼓励高效的推理模式。在SWE-bench Verified和SWE-bench Multilingual上的大量实验表明,与最先进的编码智能体和强化学习方法相比,HERO在解决率和令牌效率之间实现了有利的权衡。

英文摘要

Recently, coding agents have emerged as a dominant paradigm for real-world software engineering (SWE) scenarios, which solve complex tasks through multi-turn interactions with development environments. However, frequent interactions with environments would inevitably introduce substantial token overhead, leading to high usage costs and latency. Although recent studies have explored reducing token usage by context manipulation and interaction limits at inference time, these approaches focus on improving token efficiency while overlooking the risk of discarding task-relevant information, thus struggling to balance the trade-off between resolution rate and token efficiency. In this paper, we study a more general paradigm without suffering from the limitation, i.e., training token-efficient coding agents with promising resolution performance, which is a highly-practical yet less-explored problem. To this end, we reveal two core observations in SWE scenarios: i) Efficiency Variation: successful resolution could be achieved with fewer tokens; ii) Entropy Correlation: unproductive behaviors are associated with turn-level entropy. Motivated by observations, we propose a novel reinforcement learning framework, dubbed HERO. Specifically, HERO prioritizes task resolution over token efficiency during policy optimization and encourages efficient reasoning patterns at both trajectory and turn levels. Extensive experiments on SWE-bench Verified and SWE-bench Multilingual demonstrate that HERO achieves a favorable trade-off between resolution rate and token efficiency compared with state-of-the-art coding agents and reinforcement learning methods.

发表机构

  • Sichuan University(四川大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑