代码检测器有半衰期:LLM生成代码检测中的过时与度量幻觉
Code Detectors Have a Half-Life: Obsolescence and Metric Illusions in LLM-Generated Code Detection
浏览论文内容
中文总结 AI 辅助
本研究提出检测器半衰期概念,评估八个通用LLM评判器和三个专用检测器在七种生成器代码上的表现,发现通用LLM评判器无需训练即可超越专用检测器,但可靠性受模型、生成器和提示策略影响。
中文摘要 AI 辅助
随着代码生成模型的演进,代码检测器可能变得过时:在一个模型代际上验证过的检测器可能无法迁移到下一代。我们将这种有限的有效使用期称为检测器半衰期。我们评估了八个通用LLM评判器和三个专用检测器,对人工编写的代码以及七个生成器在C++、Java和Python中生成的代码进行了测试。我们的结果揭示了两个问题。首先,性能在不同生成器和提示策略之间差异显著,这表明一些检测器依赖于生成器特定的模式,而非代码来源的普遍证据。其次,准确率可能掩盖严重的预测偏差。DetectCodeGPT和GPT-Sniffer在所有生成器上达到了0.50的准确率,但F1分数为0.00,因为它们几乎将每个样本都分类为AI生成。然而,通用LLM评判器取得了更强的准确率和F1分数。我们的结果表明,通用LLM是很有前景的无需训练的代码来源评判器,并且可以超越专用检测器。然而,它们的可靠性取决于评判器模型、代码生成器和提示策略。因此,我们建议在多个生成器上评估LLM评判器,并报告宏F1以及类别特定的精确率和召回率。
英文摘要
Code detectors can become obsolete as code-generating models evolve: a detector validated on one generation of models may not transfer to the next. We call this limited useful life a detector half-life. We evaluate eight general-purpose LLM judges and three dedicated detectors on human-written code and code produced by seven generators across C++, Java, and Python. Our results reveal two problems. First, performance varies considerably across generators and prompting strategies, suggesting that some detectors rely on generator-specific patterns rather than general evidence of code provenance. Second, accuracy can conceal severe prediction bias. DetectCodeGPT and GPT-Sniffer achieved an accuracy of 0.50 but an F1 score of 0.00 across all generators because they classified almost every sample as AI-generated. However, general-purpose LLM judges achieved stronger accuracy and F1 scores. Our results show that general-purpose LLMs are promising training-free judges of code provenance and can outperform dedicated detectors. However, their reliability depends on the judge model, the code generator, and the prompting strategy. We therefore recommend evaluating LLM judges across multiple generators and reporting macro-F1 alongside class-specific precision and recall.
发表机构
- University of Luxembourg(卢森堡大学)
机构由 AI 辅助整理,请以论文原文为准。