AI 中文总结
研究针对编码代理长上下文修剪问题,提出SWE-Pruner Pro,利用代理内部表示直接修剪工具输出。在多基准测试中节省令牌,保持任务质量,在MiMo-V2-Flash上提升解决率和准确率。
AI 中文摘要
修剪编码代理的长上下文一直是高效上下文管理的关键技术。虽然现有的上下文修剪方法(如SWE-Pruner)通过附加单独的代码分类器来实现,但我们发现代理在读取工具输出时本身会编码表示代码上下文相关性的内部表示。基于这一发现,我们提出了SWE-Pruner Pro,它直接在代理内部修剪工具输出。具体来说,一个小的头部将代理自己的内部表示转换为每行的保留或修剪标签,同时有一个与每个工具输出行数相关的长度感知嵌入。在两个开放权重主干和四个多轮基准测试中,SWE-Pruner Pro在保持任务质量的同时,最多可节省39%的提示和完成令牌,且推理开销有限。值得注意的是,在MiMo-V2-Flash上,SWE-Pruner Pro还将SWE-Bench验证解决率提高了3.8%,将长上下文乌龙准确率提高了2.2个百分点。
英文摘要
Pruning long context for coding agents has been a vital technology for efficient context management. While existing context pruning methods such as SWE-Pruner realize this by attaching a separate code classifier, we find the agent itself encodes internal representations indicating the relevance of code context when reading tool output. Based on this finding, we propose SWE-Pruner Pro, which prunes tool outputs directly inside the agent. Concretely, a small head turns the agent's own internal representations into a keep-or-prune label for each line, with a length-aware embedding keyed to each tool output's line count. Across two open-weight backbones and four multi-turn benchmarks, SWE-Pruner Pro saves up to 39% of prompt and completion tokens while preserving task quality, with bounded inference overhead. Notably, on MiMo-V2-Flash SWE-Pruner Pro additionally raises the SWE-Bench Verified resolve rate by +3.8% and the long-context Oolong accuracy by +2.2 points.
CommentsProject page: https://github.com/Ayanami1314/swe-pruner-pro