arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

风格而非自我:表面线索解释大语言模型的零样本代码归属

Style, Not Self: Surface Cues Explain Zero-Shot Code Attribution by Large Language Models

Ehsan Barkhordar, Surendrabikram Thapa

arXiv 2609.30048首次发表:更新:

发表机构

Koç University; Virginia Tech(科奇大学; 弗吉尼亚理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究测试大语言模型零样本代码归属能力,发现表面线索(如代码长度)而非自我识别驱动其判断,建议采用平衡准确率等稳健评估指标。

AI 中文摘要

如果语言模型能够识别自己编写的代码,它可能会在作为评判者时偏爱该代码,并且一个模型监控另一个模型的实例可能会串通。我们在当前商业模型上以零样本方式测试了这一点。五个大语言模型生成MBPP、HumanEval和DS-1000的解决方案,另外七个仅生成MBPP的解决方案,模型在四个任务中充当评估者:从一对中挑选自己的解决方案,判断单个解决方案是否是自己编写的,识别两个解决方案中哪一个是由指定模型编写的,以及盲评质量。在单解决方案任务中,所有15个模型-基准组合的平衡准确率为49-58%,而原始准确率(38-67%)主要反映模型认领作者身份的倾向。在两两比较任务中,14个评估者-对手组合的准确率与评估者解决方案更长的频率之间的相关性为r=0.93。对指定模型的归属在某些组合上成功,而在其他组合上持续反转。一种基于规则的归一化方法去除文档字符串、注释、类型提示和局部名称,保留了Pass@1,并使十二个重新测试结果中的十个处于偶然水平;另外两个遵循其留下的长度差异,尽管训练的分类器仍能分离大多数归一化后的组合。Claude Haiku的自我偏好也消失了。我们建议报告平衡准确率、启发式基线和标签一致性。

英文摘要

If a language model can recognize code it wrote, it may favor that code as a judge, and instances of one model monitoring each other could collude. We test this zero-shot on current commercial models. Five LLMs generate solutions to MBPP, HumanEval, and DS-1000, seven more to MBPP, and models act as evaluators in four tasks: picking their own solution from a pair, judging whether a single solution is their own, identifying which of two solutions a named model wrote, and judging quality blind. In the single-solution task, balanced accuracy is 49-58% for all 15 model-benchmark combinations, while raw accuracy (38-67%) mostly reflects how readily a model claims authorship. In the pairwise task, accuracy across 14 evaluator-opponent combinations correlates at r=0.93 with how often the evaluator's solution is longer. Attribution to a named model succeeds on some pairs and is consistently inverted on others. A rule-based normalization that strips docstrings, comments, type hints, and local names preserves Pass@1 and leaves ten of twelve re-tested results at chance; the other two follow a length difference it leaves, although a trained classifier still separates most normalized pairs. Claude Haiku's self-preference also disappears. We recommend reporting balanced accuracy, heuristic baselines, and label consistency.

Comments18 pages, 1 figure. Code and data: https://github.com/ebarkhordar/llm-collusion

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑