arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI生成诗歌的类人特性表征:一项零样本分类研究

Characterizing Human-Likeness in AI Generated Poetry: A Zero-shot Classification Study

A. N. Biswas, T. Tabassum, A. A. Shohid, R. M. Mou, A. A. Esha, F. Sadeque, A. Ahmed

arXiv 2607.26221首次发表:更新:

发表机构

BRAC University(BRAC大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出零样本检测流水线,结合人类与AI诗歌数据集,推导影响诗歌分类及误分类的属性,为诗歌可区分性主张提供证据,同时减少训练需求并强化GenAI检测流水线。

AI 中文摘要

随着AI技术的发展,生成式AI(GenAI)文本与人类撰写的文本已几乎难以区分,AI聊天机器人的全球标准化也加剧了学术不端问题。现有研究表明,即便未经过任何修改,GenAI诗歌仍是最难区分的类型,因此现代检测器自然认为GenAI诗歌具有类人性。但这类研究的客观性需通过现代检测工具验证,而诗歌的主观性以及现代大语言模型(LLMs)架构的黑箱特性,使得验证工作相当复杂。因此,本研究的主要目标是推导对人类和AI诗歌的分类及误分类有贡献的英文诗歌属性,并为诗歌可区分性的主张提供佐证或矛盾证据。为完成此类表征,我们提出了一种零样本检测流水线,该流水线包含由人类和AI诗歌组成的数据集,用于验证人类与AI创作的可区分性,并提取上述关键属性以实现准确分类。提取此类属性有两方面益处:其一,它减少了所需的训练幅度,因为仅需对基于误分类属性的诗歌进行训练和微调;其二,它为GenAI检测难题提供了关键见解,以强化现代检测流水线。

英文摘要

With the advancement of AI technologies, Generative AI (GenAI) and human written text have become nearly indistinguishable. Additionally, the global standardization of AI chatbots made academic malpractice more frequent. Furthermore, existing research indicates GenAI poems are the most difficult to distinguish even without any modification thus, GenAI poems are naturally deemed human-like by modern detectors. However, the objectivity of such dissertations needs to be verified against modern detection tools but the subjectivity of poetry and the black-box nature of the modern LLMs (Large Language Models) architectures made verification of such work quite complicated. Hence, the main objective of the research is to deduce the attributes of English poetry that contribute classification and misclassification of both human and AI poems and provide corroborating or contradicting evidence to the poetry distinguishability claim. For such characterizations, we propose a Zero-shot detection pipeline with a dataset consisting of both human and AI poems to verify the distinguishability of human and AI creation and extract the aforementioned crucial attributes for accurate classification. Extraction of such attributes provides benefits in two ways: firstly, it reduces the margin of training needed as only the poems based on misclassifying attributes need to be trained and fine tuned and finally provides a critical insight to the GenAI detection dilemma to strengthen the modern detection pipelines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑