arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.18550cs.SE

错误修复中的语义漂移:行为信号如何从报告传播到测试和补丁

Semantic Drift in Bug Resolution: How Behavioral Signals Propagate from Reports to Tests and Patches

Wendkûuni C. Ouédraogo, Yinghua Li, Xueqi Dang, Paweł Borsukiewicz, Liang Xiao, Lingfeng Bao, Anil Koyuncu, Jacques Klein, David Lo, Tegawendé F. Bissyandé

首次发表
浏览论文内容

中文总结 AI 辅助

研究错误修复中行为信号跨工件传播的量化问题,核心方法是引入Desc2Fix框架,通过结构化行为锚等实现语义对齐,主要贡献是证明行为对齐可测且非相似性可比,能为测试和补丁相关工作提供可重现信号,助力多项错误修复任务。

中文摘要 AI 辅助

错误修复是一个跨工件的过程,自然语言报告必须揭示可由测试重现并由补丁纠正的可操作行为线索。但这些信号在工件间的保留程度很大程度上未被量化。我们引入Desc2Fix框架来衡量错误报告、触发测试和开发者编写的修复之间的语义对齐。通过结构化行为锚、确定性相似性度量和基于大语言模型的判断来实现对齐。我们使用GPT-4o和DeepSeek-Chat分析了来自Defects4J和SWT-Bench的2857个报告-测试-补丁三元组。大语言模型能可靠提取结构化信号,但对齐对表示很敏感。两个模型相对于人类都表现出系统性乐观,激励了偏差感知评估。结果表明行为对齐可测量但不能简化为相似性,结构化锚与基于嵌入的代理相结合可为测试和候选补丁的排名与过滤提供可重现信号,Desc2Fix能实现更可靠的测试生成等。

英文摘要

Desc2Fix is a framework for measuring semantic alignment between bug reports, triggering tests, and developer-written fixes. Alignment is operationalized through structured behavioral anchors (e.g., reproduction steps, API/exception cues, expected vs. actual behavior), deterministic similarity metrics (ROUGE, SBERT, CodeBERT, OpenAI embeddings), and LLM-based judgments grounded in coverage, correctness, and specificity. Our analysis covers 2,857 report-test-patch triplets from Defects4J and SWT-Bench using two widely adopted instruction-tuned LLMs from distinct model families. LLMs reliably extract structured signals (up to 90% completeness) and exhibit strong cross-model consistency, yielding a stable semantic input contract for downstream reasoning. However, alignment is highly representation-sensitive: lexical similarity alone is insufficient, full diffs provide the most stable basis for judging report-patch correspondence, and structured summaries trade surface overlap for stronger correspondence at the level of individual actions and entities. Across more than 182,000 LLM-based alignment judgments, both models exhibit systematic optimism relative to humans (1-2 points on 5-point scales) and only modest rank agreement, motivating bias-aware evaluation. Behavioral alignment is measurable but not reducible to similarity, and structured anchors combined with embedding-based proxies provide reproducible signals for ranking and filtering tests and patches. Desc2Fix could support more reliable test generation, fault localization, patch ranking, and bug report authoring.

↑