arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05917cs.SE

跳出自修复陷阱:通过双上下文感知改进测试预言生成

Escaping the Self-Repair Trap: Improving Test Oracle Generation via Dual-Context Awareness

Kefan Li, Hongyue Yu, Yuan Yuan

首次发表
浏览论文内容

中文总结 AI 辅助

针对迭代自修复引发的自修复陷阱问题,提出非迭代高效框架DCAware,结合结构化静态上下文与动态状态,在低计算成本下提升测试预言的故障检测效果

中文摘要 AI 辅助

大语言模型(LLMs)在回归预言补全任务中展现出强大潜力,该任务给定测试前缀,将当前程序版本视为预期行为。近期方法愈发依赖迭代自修复与执行反馈,但优化执行成功度未必能生成强故障检测能力的预言。这种在修复类方法中广泛采用的目标仅为代理指标,可能与预言生成的真实目标不一致,这种不一致会使修复过程产生偏差,引发反馈驱动的退化,我们将其称为“自修复陷阱”,即迭代修复会逐步推动模型趋向于更易满足但检测故障能力更弱的断言。为解决该问题,我们提出DCAware,这是一种计算高效的非迭代框架,优先考虑高信噪比的上下文基础而非多轮修复。DCAware将结构化静态上下文与选择性检索的动态状态相集成,无需迭代反馈循环即可实现精准且鲁棒的预言生成。基于执行测试与变异测试的大量实验表明,DCAware在保持高执行成功度的同时,持续提升故障检测有效性,且计算成本显著低于现有方法。我们的结果表明,在所研究的回归预言场景中,提升上下文质量比增加迭代修复复杂度更有效。

英文摘要

Large Language Models (LLMs) have shown strong potential for regression-oracle completion, where a test prefix is given and the current program version is treated as expected behavior. Recent approaches increasingly rely on iterative self-repair and execution feedback, but optimizing execution success does not necessarily yield strong fault-revealing oracles. This objective, widely adopted in repair-based methods, serves only as a proxy and may be misaligned with the true goal of oracle generation. Such misalignment biases the repair process, giving rise to a feedback-driven degeneration that we term the Self-Repair Trap, where iterative repair progressively drives models toward assertions that are easier to satisfy but less effective at detecting faults. To address this issue, we propose DCAware, a computationally efficient, non-iterative framework that prioritizes high signal-to-noise contextual grounding over multi-round repair. DCAware integrates structured static context with selectively retrieved dynamic states, enabling precise and robust oracle generation without iterative feedback loops. Extensive experiments based on execution and mutation testing show that DCAware consistently improves fault-revealing effectiveness while maintaining high execution success, outperforming prior methods with substantially lower computational cost. Our results suggest that improving contextual quality is more effective than adding iterative repair complexity in the studied regression-oracle setting.

补充信息

↑