RCL:用于检测企业检索增强代码生成中上下文不足的检索置信层
RCL: A Retrieval-Confidence Layer for Detecting Insufficient Context in Enterprise Retrieval-Augmented Code Generation
浏览论文内容
中文总结 AI 辅助
针对企业私有代码库中检索增强代码生成因上下文不足而失败的问题,提出检索置信层RCL,结合结构覆盖与新颖性分数在生成前检测检索不足,触发后续检索或人工审查,提升生成正确性。
中文摘要 AI 辅助
检索增强生成(RAG)在代码生成方面已在公共代码库上得到广泛研究,在这些代码库中,模型的参数化知识往往能弥补检索不完善的问题。然而,这在企业代码库中失效,因为私有API、内部框架和未记录的团队约定完全超出任何模型的预训练分布。近期关于私有库代码生成的研究表明,即使使用完美检索(oracle retrieval)也无法消除错误,而是将失败定位在API使用环节;另外,置信门控检索已被研究用于开放域问答,利用模型内部置信度。这两者都未解决在生成开始前检索本身对私有代码查询是否在结构上充分的问题。我们提出RCL(检索置信层),一种插入在检索与生成之间的轻量级模块,结合基于调用图的结构覆盖分数与估计查询对模型先验知识之外依赖程度的新颖性分数,在生成发生前检测检索不足。当置信度低于校准阈值时,RCL触发有针对性的后续检索或将输出标记供人工审查,而不是在上下文不完整的情况下静默生成。我们描述了RCL的架构,形式化了其评分函数,并提出了一种评估方法,使用通过向开源Java仓库注入合成内部API构建的私有代码基准,在无专有代码的情况下模拟企业条件。我们报告了结果(第7节),将RCL与仅基于相似性的检索在生成正确性方面进行比较。我们的立场是,检索充分性(从结构上而非从模型内部置信度评估)是在私有企业环境中构建更安全代码生成系统的一个独特且当前未被充分处理的信号。
英文摘要
Retrieval-Augmented Generation (RAG) for code generation has been studied extensively on public repositories, where a model's parametric knowledge often compensates for imperfect retrieval. This breaks down in enterprise codebases, where private APIs, internal frameworks, and undocumented team conventions fall entirely outside any model's pretraining distribution. Recent work on private-library code generation shows that even oracle (perfect) retrieval does not eliminate errors, but locates failures downstream in API usage; separately, confidence-gated retrieval has been studied for open-domain question answering using model-internal confidence. Neither addresses whether retrieval itself was structurally sufficient for a private-code query before generation begins. We introduce RCL (Retrieval-Confidence Layer), a lightweight module inserted between retrieval and generation that combines a call-graph-derived structural coverage score with a novelty score estimating a query's dependence on knowledge outside the model's prior, to detect insufficient retrieval before generation occurs. When confidence falls below a calibrated threshold, RCL triggers a targeted follow-up retrieval or labels the output for human review, rather than generating silently against incomplete context. We describe RCL's architecture, formalize its scoring functions, and propose an evaluation methodology using a private-code benchmark built by injecting synthetic internal APIs into open-source Java repositories, simulating the enterprise condition without proprietary code. We report results (Section 7) comparing RCL against similarity-only retrieval on generation correctness. Our position is that retrieval sufficiency, assessed structurally rather than from model-internal confidence, is a distinct and currently underaddressed signal for building safer code-generation systems in private, enterprise settings.