arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12489cs.AIcs.CLcs.SE

置信门控的直推式测试生成用于代码重排序

Confidence-Gated Transductive Test Generation for Code Reranking

Sungjae Lee, Youngsik Yoon, Seockbean Song, Siwei Wang, Wei Chen, Jungseul Ok

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM生成代码的测试用例构建难题,提出置信门控的直推式测试生成方法,仅在归纳置信度低时启用直推式生成,在代码重排序基准上以更低成本超越基线,实现效率与效果的平衡。

中文摘要 AI 辅助

测试用例合成对于评估和排序由大型语言模型(LLMs)生成的程序至关重要。然而,构建高质量的测试用例仍然具有挑战性,因为可靠的预期输出往往难以获得。我们提出了置信门控的直推式测试生成(CoTT),该方法首先使用高效的归纳过程,仅在归纳置信度较低时才调用直推式生成。这种自适应设计提高了输出可靠性,同时仅在需要时分配额外的计算资源。在代码重排序基准测试中,CoTT在报告的各项指标上优于先前的基线方法,同时相对于对每个输入都应用直推式生成,降低了成本。这些结果表明,基于置信度的测试时计算分配在仅使用单个高效LLM的情况下,提供了有利的效率-效果权衡。

英文摘要

Test case synthesis is crucial for evaluating and ranking programs generated by large language models (LLMs). However, constructing high-quality test cases remains challenging because reliable expected outputs are often difficult to obtain. We propose Confidence-Gated Transductive Test Generation (CoTT), which first uses an efficient inductive procedure and invokes transductive generation only when inductive confidence is low. This adaptive design improves output reliability while allocating extra computation only when needed. On code reranking benchmarks, CoTT outperforms prior baselines across the reported metrics while reducing cost relative to applying transductive generation to every input. These results show that confidence-based allocation of test-time computation provides a favorable efficiency-effectiveness trade-off with a single efficient LLM.

发表机构

  • POSTECH(浦项科技大学)
  • Microsoft Research Asia(微软亚洲研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑