arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17937cs.SE

在长上下文情况下代理技能如何失效:代码审计中的白盒研究

When and How Context Rot Appears in Coding Agents: A White-Box Study of Agent Skills in Code Auditing

Yue Xue

AI总结:

研究长上下文下代码审计中代理技能失效问题,通过改变上下文分类故障位置,对比不同条件下Codex等运行情况,发现长上下文影响大,不同任务表现有别,外部检查表效果好,编码代理支架有帮助,给出故障分类和实证案例。

AI中文摘要:

代理技能包包含程序指令和检查以供通用代理使用,但加载技能并不能保证在长时间使用工具的轨迹中每个要求都保持有效。我们在源自生产的白盒代码审计工作流程中研究此问题。固定任务和24个工件检查,改变周围上下文并对故障首次可见的位置进行分类:需求丢失、编辑漂移、检查失败或非代理评估器/运行时故障。Codex with gpt - 5.4 - mini在10991字符的干净上下文中10次运行通过8次,但在299140字符的相关上下文和等长的不相关上下文中仅通过3次。第二个任务在所有干净和长时间运行中都通过。详细的外部检查表10次运行通过10次,通用自我检查为5次。编码代理支架可能有帮助但不能消除故障。我们提供了白盒代码审计的有界故障分类和实证案例研究。

英文摘要:

Agent Skills package procedural instructions and checks for use by general-purpose agents, but loading a skill does not guarantee that every requirement remains active throughout a long tool-using trajectory. We study this problem in a production-derived, white-box code-audit workflow. Holding the task and 24 artifact checks fixed, we vary the surrounding context and classify where failures first become visible: lost requirements, editing drift, failed checking, or non-agent evaluator/runtime failures. Codex with gpt-5.4-mini passes 8/10 runs in a 10,991-character clean context but only 3/10 in both a 299,140-character relevant context and an equal-length irrelevant context. This 50-percentage-point difference is large but remains trend-level under two-sided Fisher tests (p = 0.0698). Requirement coverage nevertheless stays above 92% in both long conditions, showing that a few omissions can invalidate an otherwise complete artifact. A second task passes all clean and long runs, so the evidence does not support a universal context-length threshold. A detailed external checklist passes 10/10 runs, compared with 5/10 for a generic self-check (p = 0.0325). Coding-agent scaffolds may help by selecting a smaller working set, but they do not eliminate failures. We do not introduce context rot or a new general monitoring method; we provide a bounded failure classification and empirical case study for white-box code auditing.

↑