面向1C:Enterprise的自然语言代码检索:一个开放基准及高效双编码器
Natural Language Code Retrieval for 1C:Enterprise: An Open Benchmark and Efficient Bi-Encoder
浏览论文内容
中文总结 AI 辅助
针对1C:Enterprise生态系统缺乏开放代码检索基准与专用模型的问题,本文构建了含3413对的开放基准,基于784057个合成三元组微调专用双编码器,在多指标上优于基线,且256维截断可大幅降本。
中文摘要 AI 辅助
自然语言代码检索是计算机科学领域快速发展的任务。然而,1C:Enterprise生态系统结合了俄语语法与高度领域特定的术语,目前几乎不存在针对该生态系统的开放数据集和专用模型。本文提出了一套用于1C代码检索的完整流程:包含3413个经PII脱敏处理的真实查询-代码对的开放基准、可复现的评估工具,以及一个专用双编码器。为解决标注数据稀缺问题,我们使用google/gemma-4-26B-A4B-it从公共代码仓库生成的784057个合成三元组进行微调,采用了嵌套表示学习(Matryoshka Representation Learning,MRL)和隐私感知分词器。由于基准子集规模不同,我们报告了平衡子集宏平均、查询加权微平均以及仅论坛子集的结果。我们的模型在平衡子集宏平均nDCG@10指标上达到0.5992,微平均为0.5044,论坛子集为0.4617,而基线架构的宏平均为0.4932,google/embeddinggemma-300m为0.5404。移除经保守的精确/13-gram重叠审计标记的每个基准示例后,平衡子集宏平均提升至0.6011(微平均为0.5010),表明检测到的训练-基准重叠并不能解释主要结果。将MRL截断至256维可保留99.9%的检索质量,同时将密集索引存储和精确相似度计算的资源消耗降低至原来的三分之一。
英文摘要
Natural language code retrieval is a rapidly evolving task in computer science. However, the 1C:Enterprise ecosystem combines Russian syntax with highly domain-specific terminology, for which open datasets and specialized models have been virtually non-existent. We present a comprehensive pipeline for 1C code retrieval: an open benchmark of 3,413 real-world, PII-scrubbed query-code pairs, a reproducible evaluation harness, and a specialized bi-encoder. To overcome scarce labeled data, we fine-tune on 784,057 synthetic triplets generated by google/gemma-4-26B-A4B-it from public code repositories, using Matryoshka Representation Learning (MRL) and a privacy-aware tokenizer. Because the benchmark subsets differ in size, we report balanced-subset macro, query-weighted micro, and forum-only results. Our model reaches 0.5992 balanced macro nDCG@10, 0.5044 micro, and 0.4617 on forum, versus 0.4932 macro for the baseline architecture and 0.5404 for google/embeddinggemma-300m. Removing every benchmark example flagged by the conservative exact/13-gram overlap audit leaves 0.6011 balanced macro (0.5010 micro), indicating that detected train-benchmark overlap does not explain the headline result. MRL truncation to 256 dimensions preserves 99.9% of retrieval quality while reducing dense-index storage and exact similarity arithmetic by a factor of three.