arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17139cs.SEcs.CL

SynH-Rank:通过多样化数据合成和分层排序训练实现质量感知代码搜索

SynH-Rank: Quality-Aware Code Search via Diverse Data Synthesis and Hierarchical Ranking Training

Keyu Liang, Haoye Wang, Yanfu Yan, Zhiyuan Wan, Zhongxin Liu

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对代码搜索中重排器忽视非功能质量的问题,提出SynH-Rank框架,结合大语言模型数据合成与分层排序训练,采用三级标注方案及新指标,有效提升了代码搜索质量,在多方面取得优于传统方法的性能。

中文摘要 AI 辅助

代码搜索通过实现高效的代码重用提高开发者生产力。当前代码搜索系统常采用检索-重排管道,重排器专注于建模查询与代码间的语义相关性,却忽略了执行速度、内存使用和可维护性等关键非功能质量。实现质量感知代码搜索面临两大挑战:高质量标注数据集稀缺及标准对比学习目标的局限性。为解决这些挑战,我们提出SynH-Rank,一种质量感知代码重排框架,它将基于大语言模型的多样化数据合成与分层排序训练相结合。SynH-Rank采用三级标注方案明确建模层次结构:高质量相关>低质量相关>不相关。此外,我们引入一个包含4209对的新基准和两个新指标:用于评估高质量代码优先级的质量偏好准确率(QPA)和用于评估复杂条件下性能的多条件准确率(MCA)。实验结果表明,SynH-Rank比主干模型将QPA提高了20.15%,比仅基于相关性的标准对比训练性能提升了15.80%,同时增强了传统相关性指标和多条件泛化能力。

英文摘要

Code search enhances developer productivity by enabling efficient code reuse. Current code search systems often use a retrieve-then-rerank pipeline, where rerankers focus on modeling semantic relevance between queries and code. However, these rerankers overlook critical non-functional qualities like execution speed, memory usage, and maintainability, which are essential for practical software development. Studies reveal developers expect results to maintain high coding standards and satisfy specific needs, such as resource optimization, highlighting the importance of quality-aware code search. Achieving quality-aware code search faces two major challenges: the scarcity of quality-annotated datasets for effective training and the limitations of standard contrastive learning objectives, which fail to capture the ordinal relationships among high-quality, low-quality, and irrelevant code. Although contrastive learning excels in distinguishing relevant from irrelevant code, its binary objective does not support nuanced quality distinctions. To address these challenges, we propose SynH-Rank, a quality-aware code reranking framework that combines LLM-driven diverse data synthesis with hierarchical ranking training. SynH-Rank employs a three-level labeling scheme to explicitly model the hierarchy: high-quality relevant > low-quality relevant > irrelevant. Additionally, we introduce a new benchmark with 4,209 pairs and two novel metrics: Quality Preference Accuracy (QPA) for assessing prioritization of high-quality code and Multi-Condition Accuracy (MCA) for evaluating performance under complex constraints. Experimental results show SynH-Rank improves QPA by 20.15\% over backbone models and outperforms standard relevance-only contrastive training by 15.80\%, while simultaneously enhancing traditional relevance metrics and multi-condition generalizability.

发表机构

  • College of Computer Science and Technology(计算机科学与技术学院)
  • The State Key Laboratory of Blockchain and Data Security(区块链与数据安全国家重点实验室)
  • Zhejiang University(浙江大学)
  • Hangzhou City University(杭州市城市大学)
  • Zhejiang University and Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security(浙江大学和杭州高新区(滨江)区块链与数据安全研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑