源代码作者身份归属无法从竞赛场景泛化到课堂场景
Source Code Authorship Attribution Does Not Generalize from Competitions to Classrooms
浏览论文内容
中文总结 AI 辅助
本研究发现,针对Google Code Jam竞赛数据微调的CodeBERT等模型,在课程作业数据集上的作者身份归属性能远低于随机基线,竞赛场景的基准无法泛化到教育场景,需针对性验证。
中文摘要 AI 辅助
源代码作者身份归属旨在通过编程风格识别程序片段的作者。我们对预训练Transformer模型CodeBERT进行微调,使用三类数据:Kaggle公开仓库中的Google Code Jam(GCJ)公开提交代码、整理后的GCJ存档,以及某理工高校收集的课程作业数据集。针对多轮GCJ数据,CodeBERT在10位作者时Top-1准确率达92.6%,在1000位作者时仍保持70.7%的Top-1准确率(Top-10准确率为88.2%)。但在被考察的课程作业数据集上,同一流程的性能等于或低于对应随机基线:在含690位作者的闭卷作业数据集上Top-1准确率为0.2%,在含812位作者的开卷作业数据集上Top-1准确率为0.06%。配套的多模型基准在其他数据集配置中提供了一致的跨模型证据,表明观测到的性能差距并非CodeBERT特有,而是在被评估的所有模型家族中均存在。我们分析了可能解释该差距的数据集与任务属性,并指出:基于GCJ的基准若未在目标课程作业场景中验证,会高估作者身份归属在教育场景中的实际适用性。
英文摘要
Source code authorship attribution aims to identify the author of a program fragment from its writing style. We fine-tune the pre-trained transformer CodeBERT on three sources of data: publicly available Google Code Jam (GCJ) submissions from an open Kaggle repository, a curated GCJ archive, and institutional coursework datasets collected at a technical university. On multi-round GCJ data, CodeBERT reaches 92.6% Top-1 accuracy for 10 authors and retains 70.7% Top-1 (88.2% Top-10) for 1000 authors. On the examined coursework datasets, the same pipeline performs at or below the corresponding chance baselines: 0.2% Top-1 on a closed-assignment dataset of 690 authors and 0.06% Top-1 on open-ended assignments evaluated over 812 authors. A companion multi-model benchmark provides consistent cross-model evidence across additional dataset configurations, indicating that the observed performance gap is not specific to CodeBERT and persists across the evaluated model families. We analyze dataset and task properties that plausibly explain this gap and argue that GCJ-based benchmarks overestimate the practical applicability of authorship attribution in educational settings unless they are validated on the target coursework context.