arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

小型Transformer中用于对比代码表示学习的合成语义监督:一项实证研究

Synthetic Semantic Supervision for Contrastive Code Representation Learning in Small Transformers: An Empirical Study

Kenneth Paulsen, Florian Tambon, Mike Papadakis, Shin Yoo

arXiv 2609.03702首次发表:更新:

发表机构

University of Luxembourg; School of Computing, KAIST(卢森堡大学; 韩国科学技术院计算机学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对小型Transformer的代码表示学习,提出用合成语义监督替代人工文档字符串或执行轨迹,经8项任务验证,该方法可提升性能且具可扩展性。

AI 中文摘要

通用代码嵌入为代码搜索、分类和检索提供支撑工具。面向代码的紧凑Transformer编码器通常依赖人工编写的文档字符串(耗时且不一致)或挖掘的结构信号(如执行轨迹,特定于设置且收集成本高)。本研究对一种替代方案展开实证研究:在训练阶段采用双编码器框架,将小型编码器与合成生成的、强调代码功能和意图的自然语言描述进行对比预训练,该自然语言描述在推理阶段被舍弃。我们在C、C++和Java的8个检索、分类及生成任务上,将该方法与基于预训练的基线、通用大语言模型(LLM)及特定嵌入模型进行基准测试。合成语义监督在8项任务中的5项上,与相同推理规模的预训练基线相比取得统计显著提升,另外2项任务表现相当;微调后,其在分类任务上可匹配或超越规模大两个数量级的零样本模型,且在匹配预训练数据时与感知执行的监督表现相当,表明其是现有代码表示范式的可扩展、有效替代方案。

英文摘要

General-purpose code embeddings power tools for code search, classification, and retrieval. Compact transformer encoders for code typically rely on either human-written docstrings (labor-intensive and inconsistent) or mined structural signals such as execution traces (setting-specific and costly to collect). We empirically study an alternative: contrastive pretraining of small encoders with synthetically generated natural-language descriptions emphasizing code functionality and intent, paired with code in a dual-encoder framework at training and discarded at inference. We benchmark this approach against pretraining-based baselines, generalist LLMs, and embedding-specific models on eight retrieval, classification, and generation tasks across C, C++, and Java. Synthetic semantic supervision yields statistically significant gains over pretraining baselines of the same inference-time size on five of eight tasks, with parity on two more; once fine-tuned, it matches or exceeds zero-shot models two orders of magnitude larger on classification, and it stays on par with execution-aware supervision at matched pretraining data, suggesting a scalable, effective alternative to existing code-representation paradigms.

CommentsAccepted in Findings EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑