arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AttnCompress:面向软件工程智能体的动态注意力引导轨迹压缩

AttnCompress: Dynamic Attention-Guided Trajectory Compression for Software Engineering Agents

Zhengran Zeng, Yixin Li, Rui Xie, Wei Ye, Shikun Zhang

arXiv 2609.08318首次发表:更新:

发表机构

Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对软件工程智能体长轨迹导致的上下文瓶颈,提出动态注意力引导压缩框架AttnCompress,通过PPL分割、注意力权重评估和滚动窗口机制,在SWE-Bench基准上以53.17%通过率超越基线并显著降本。

AI 中文摘要

从以人为中心的辅助向自主软件工程(ASE)智能体的转变,使得解决复杂的现实世界软件工程任务成为可能。然而,这些智能体试错式的特性生成了冗长的交互轨迹,在上下文窗口限制和成本方面造成了严重的瓶颈。虽然上下文压缩提供了一种潜在的补救措施,但先前的方法受限于静态剪枝策略和粒度不匹配,往往无法保留对软件工程任务至关重要的语义依赖和句法细节。为了在缩减上下文长度的同时严格保留关键任务证据,我们提出了AttnCompress,一种动态注意力引导的轨迹压缩框架。与现有方法不同,AttnCompress通过三个关键机制弥合了语义完整性与动态适应性之间的鸿沟:(1)利用困惑度(PPL)尖峰进行结构感知分割,以保留代码和日志的句法结构;(2)使用代理注意力权重进行相关性估计,以量化历史块与智能体当前推理之间的精确相关性;(3)采用动态滚动窗口,随着任务演进重新评估并召回历史上下文。在SWE-Bench-Verified和Multi-SWE-Bench上的广泛评估表明,AttnCompress实现了53.17%的通过率,超越了先前最先进的基线,同时将令牌消耗降低了21.6%,总成本降低了33.6%。该框架被证明是模型无关的,并能有效泛化到多种编程语言。

英文摘要

The transition from human-centric assistance to Autonomous Software Engineering (ASE) agents has enabled the resolution of complex real-world SE tasks. However, the trial-and-error nature of these agents generates lengthy interaction trajectories, creating severe bottlenecks in terms of context window limits and cost. While context compression offers a potential remedy, prior approaches suffer from static pruning strategies and granularity mismatches, often failing to preserve the semantic dependencies and syntactic details crucial for SE tasks. To strictly preserve critical task evidence while reducing context length, we introduce AttnCompress, a dynamic attention-guided trajectory compression framework. Unlike existing approaches, AttnCompress bridges the gap between semantic integrity and dynamic adaptability through three key mechanisms: (1) structure-aware segmentation via perplexity (PPL) spikes to preserve the syntactic structure of code and logs; (2) relevance estimation using proxy attention weights to quantify the precise relevance of historical blocks to the agent's current reasoning; and (3) a dynamic rolling window to re-evaluate and recall historical context as the task evolves. Extensive evaluation on SWE-Bench-Verified and Multi-SWE-Bench demonstrates that AttnCompress achieves a pass rate of 53.17%, outperforming prior state-of-the-art baselines while reducing token consumption by 21.6% and total costs by 33.6%. The framework proves to be model-agnostic and generalizes effectively across diverse programming languages.

Comments23 pages, 5 figures, accepted at ISSTA 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑