arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38402cs.SEcs.LG

从代码库到元凶(C2C):利用语义检索与分层强化学习缩减缺陷搜索空间

From Codebase to Culprit (C2C): Reducing the Search Space for Bugs with Semantic Retrieval and Hierarchical Reinforcement Learning

Ankur Garg, Corey Yang-Smith, Rishav Rishav, Ahmad Abdellatif, Samira Ebrahimi Kahou

首次发表
浏览论文内容

中文总结 AI 辅助

C2C框架结合语义检索与分层强化学习,在文件、函数和代码行多粒度上逐步缩减缺陷搜索空间,实现精确缺陷定位,并在真实Java和Python数据集上提升检索精度与定位准确性。

中文摘要 AI 辅助

我们引入了C2C(从代码库到元凶),一个用于精确缺陷定位的框架,它逐步在多个粒度级别(文件、函数和代码行)上缩减调试搜索空间。为了模拟开发者自然的自顶向下调试工作流,C2C在两步过程中整合了语义检索和分层强化学习(HRL)。首先,它通过使用缺陷报告文本(包括可用的堆栈跟踪信息)进行语义向量相似性搜索,对候选缺陷进行面向召回率的检索,该搜索针对嵌入数据库进行,其中嵌入通过使用CodeBERT的对比学习进行微调。基于这个缩减的搜索空间,HRL框架逐步定位缺陷,从文件推理到函数,最终到单独的代码行。与先前在单一粒度上操作的方法不同,C2C实现了多分辨率定位,同时保持跨决策的上下文一致性。在真实世界的Java和Python数据集上的实验表明,C2C提高了检索精度和定位准确性。消融研究进一步强调了层次分解、结构化学习信号和奖励塑形在推进多级缺陷定位中的贡献。

英文摘要

We introduce C2C (From Codebase to Culprit), a framework for precise bug localization that progressively reduces the debugging search space across multiple levels of granularity: files, functions, and lines of code. To mirror developer's natural top-down debugging workflows, C2C integrates semantic retrieval and Hierarchical Reinforcement Learning (HRL) in a two-stage process. First, it performs recall-oriented retrieval of buggy candidates via semantic vector similarity search using bug-report text, including available stack-trace information, against a database of embeddings, where the embeddings are fine-tuned via contrastive learning with CodeBERT. Building on this reduced search space, the HRL framework incrementally localizes bugs, reasoning from files to functions and ultimately to individual lines of code. Unlike prior approaches which operate at a single granularity, C2C enables multi-resolution localization while maintaining contextual consistency across decisions. Experiments on real-world Java and Python datasets demonstrate that C2C improves retrieval precision and localization accuracy. Ablation studies further highlight the contributions of hierarchical decomposition, structured learning signals, and reward shaping in advancing multi-level bug localization.

发表机构

  • University of Calgary(卡尔加里大学)

机构由 AI 辅助整理,请以论文原文为准。

↑