arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

区分人工与真实:评估大语言模型检测由大语言模型生成的内容

Distinguishing Artificial from Authentic: Evaluating LLMs for Detecting LLM-Generated Content

Juho Leinonen, Paul Denny

arXiv 2607.20446首次发表:更新:

发表机构

Aalto University; University of Auckland(阿尔托大学; 奥克兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究大语言模型检测自身生成内容的能力,通过真实学生回答和模型生成答案的多种变体,评估不同提示策略和输出格式下的检测性能,发现其检测高度依赖任务,提示框架等因素影响检测效果,凸显了大语言模型自我检测的潜力与局限。

AI 中文摘要

随着学生越来越多地使用大语言模型(LLMs)来生成自然语言回答和程序代码,人们越来越关注大语言模型本身是否可用于区分人工智能生成的作业和人类撰写的提交内容。本文研究了大语言模型在多种教育任务类型中检测自身生成内容的程度,包括编程练习、反思性写作和简答题。使用真实学生回答和大语言模型生成答案的多种变体,我们评估了不同提示策略和输出格式下的检测性能。我们的研究解决了三个研究问题:(1)大语言模型在不同任务领域中识别自身输出的准确程度;(2)提示设计、回答长度和任务类型等因素如何影响检测效果;(3)大语言模型生成回答的哪些特征导致检测成功或失败。我们的发现表明,基于大语言模型的检测高度依赖任务:对于编程任务和较长的反思性回答,检测更可靠,但对于简答题表现不佳,大语言模型经常将自己的输出判断为比真实学生回答更像人类。我们还发现,提示框架和回答冗长程度对反思性写作任务的可检测性有显著影响,提示的相对较小变化会显著降低检测准确性,而与编程相关的检测对提示变化更具鲁棒性。这些结果凸显了大语言模型在教育环境中自我检测的潜力和局限性,并建议在依赖大语言模型作为识别人工智能生成的学生作业的独立工具时要谨慎。

英文摘要

As large language models (LLMs) are increasingly used by students to generate natural language responses and program code, there is growing interest in whether LLMs themselves can be used to distinguish AI-generated work from human-authored submissions. In this paper, we investigate the extent to which LLMs can detect their own generated content across multiple educational task types, including programming exercises, reflective writing, and short-answer questions. Using authentic student responses and multiple variants of LLM-generated answers, we evaluate detection performance under different prompting strategies and output formats. Our study addresses three research questions: (1) how accurately LLMs can identify their own outputs across task domains, (2) how detection effectiveness is influenced by factors such as prompt design, response length, and task type, and (3) what characteristics of LLM-generated responses contribute to successful or failed detection. Our findings show that LLM-based detection is highly task-dependent: detection is substantially more reliable for programming tasks and longer reflective responses, but performs poorly for short-answer questions, where LLMs frequently judge their own outputs as more human-like than authentic student responses. We further find that prompt framing and response verbosity have a pronounced effect on detectability in reflective writing tasks, with relatively minor prompt variations significantly reducing detection accuracy, while programming-related detection is more robust to prompt changes. Together, these results highlight both the potential and the limitations of LLM self-detection in educational settings and suggest caution in relying on LLMs as standalone tools for identifying AI-generated student work.

Comments8 pages, 5 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑