AI 中文总结
MergeSE是一款开源工具,可对HuggingFace编码器检查点进行无训练合并,支持软件工程多类分类任务,能解决代码模型分布偏移导致的性能下降问题,合并效率高且性能优异。
AI 中文摘要
经过微调的代码模型通常表现为领域专家,且在分布偏移下性能会急剧下降:在我们的克隆检测场景中,在同语言克隆上训练的模型在跨语言克隆上的F1值下降71%,而多任务训练模型在未见过的AI生成克隆上的F1值降至0.151。我们的配套研究表明,事后模型合并可解决这种性能碎片化问题,无需训练数据即可达到多任务性能的93%,且对未见过的克隆类型的泛化能力提升4倍。然而,目前尚无实用工具能让软件工程(SE)研究人员检查检查点兼容性、合并专家模型、在SE基准上验证结果并导出用于部署的模型。我们提出的MergeSE是一款用于HuggingFace编码器检查点的无训练模型合并的开源命令行界面(CLI)和网页工具。尽管MergeSE由克隆检测中的分布外(OOD)泛化需求驱动,但它通过内置的9种任务类型注册表(包括漏洞检测、缺陷预测和代码异味检测),更广泛地支持SE分类工作流。MergeSE提供5种操作:tasks(任务)、inspect(检查)、merge(合并)、evaluate(评估)和export(导出)。它支持5种合并算法,包括TIES、DARE-TIES、Wudi、PCB和平均法;可检测跨任务分类头不匹配;产生可设置种子的确定性输出;并包含用于冒烟测试复现的捆绑基准样本。在CPU上,两个1.24亿参数检查点的完整合并耗时不到5秒。端到端验证确认,MergeSE生成的检查点与参考实现匹配,并能从领域专家模型中恢复跨领域性能。该工具在线可访问,开发仓库位于此URL。
英文摘要
Fine-tuned code models often behave as domain specialists and can degrade sharply under distribution shift: in our clone-detection setting, a model trained on same-language clones drops 71\% F1 on cross-language clones, while multi-task training falls to 0.151 F1 on unseen AI-generated clones. Our companion study shows that post-hoc model merging can address this fragmentation, achieving 93\% of multi-task performance without training data while generalizing 4$\times$ better to unseen clone types. However, no practical tool exists that lets SE researchers diagnose checkpoint compatibility, merge specialists, validate results on SE benchmarks, and export models for deployment. We present \textbf{MergeSE}, an open-source CLI and web tool for training-free model merging of HuggingFace encoder checkpoints. While motivated by OOD generalization in clone detection, MergeSE supports SE classification workflows more broadly through a built-in registry of nine task types, including vulnerability detection, defect prediction, and code-smell detection. MergeSE provides five operations: \textit{tasks}, \textit{inspect}, \textit{merge}, \textit{evaluate}, and \textit{export}. It supports five merging algorithms, including TIES, DARE-TIES, Wudi, PCB, and averaging; detects cross-task classification-head mismatches; produces seedable deterministic outputs; and includes bundled benchmark samples for smoke-test reproduction. A full merge of two 124M-parameter checkpoints completes in under 5 seconds on CPU. End-to-end validation confirms that MergeSE-produced checkpoints match reference implementations and recover cross-domain performance from domain-specific specialists. The tool is available online at https://mergese.usask.ca, and the development repository is at https://github.com/srlabUsask/MergeSE.
CommentsAccepted at ASE 2026