arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EffiHolmes:基于差分分析的仓库级时间低效修复定位

EffiHolmes: Differential Profiling-Guided Repository Level Time Inefficiency Fix Localization

Haowen Yang, Yun Peng, Zishuo Ding

arXiv 2608.03558首次发表:更新:

AI 中文总结

EffiHolmes是基于LLM的仓库级时间低效修复定位框架,通过差分分析等技术提升定位精度,在RepoEffi-Bench上优于多种基线方法,且跨模型规模稳健。

AI 中文摘要

大型软件系统常存在时间低效问题,这类问题虽功能正确却会导致执行时间过长。定位其修复位置十分困难,因为与功能缺陷不同,它们既不会产生测试失败,也不会提供堆栈跟踪线索,使得传统及近期基于大语言模型(LLM)的缺陷定位方法均不适用。运行时分析提供了替代证据,但在仓库级设置中面临三项挑战:单次运行的分析无法可靠区分低效热点与执行噪声;现有分析工具难以从大量后台执行中提取相关执行路径;观测到的热点与实际修复位置间存在语义差距。我们提出EffiHolmes,一种用于仓库级时间低效修复定位的LLM框架。EffiHolmes采用默认和缩放工作负载下的差分分析来识别低效热点,提取连接这些热点与报告低效函数的紧凑执行路径,并利用领域引导的LLM推理定位潜在的低效逻辑。我们还推出RepoEffi-Bench,首个用于仓库级低效定位的基准,包含从流行Python仓库收集的140个高质量问题。实验表明,EffiHolmes始终优于最先进的基于检索、智能体和分析的基线方法,使用GPT-5.1时文件级Acc@3提升4.29个百分点,使用qwen3-4b时函数级Acc@5提升15.00个百分点,且在不同模型规模下均保持稳健性。

英文摘要

Large software systems often suffer from time inefficiencies that cause excessive execution time despite functional correctness. Localizing their fix locations is difficult because, unlike functional bugs, they produce neither test failures nor stack-trace clues, making traditional and recent LLM-based fault localization methods unsuitable. Runtime profiling provides alternative evidence but faces three challenges in repository-level settings: single-run profiling cannot reliably distinguish inefficiency hotspots from execution noise; existing profilers struggle to extract relevant execution paths from extensive background execution; and a semantic gap remains between observed hotspots and actual fix locations. We propose EffiHolmes, an LLM-based framework for repository-level time inefficiency fix localization. EffiHolmes uses differential profiling under default and scaled workloads to identify inefficiency hotspots, extracts compact execution paths connecting these hotspots to the reported inefficient function, and employs domain-guided LLM reasoning to locate the underlying inefficiency logic. We also introduce RepoEffi-Bench, the first benchmark for repository-level inefficiency localization, containing 140 high-quality issues collected from popular Python repositories. Experiments show that EffiHolmes consistently outperforms state-of-the-art retrieval-, agent-, and profiling-based baselines, improving file-level Acc@3 by 4.29 percentage points with GPT-5.1 and function-level Acc@5 by 15.00 percentage points with qwen3-4b. It also remains robust across model capacities.

CommentsAccepted at the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE 2026). 13 pages, 3 figures, and 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑