发表机构
University of Maryland; Google DeepMind; Simon Fraser University(马里兰大学; 谷歌DeepMind; 西蒙弗雷泽大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
IdeaLens通过将文档转化为大纲来检测创意来源,区分人类与AI创意,在多种基准上保持高性能,并揭示系统性差异。
AI 中文摘要
虽然现代AI检测器能识别谁写了文字,但新兴的AI使用政策越来越取决于另一个问题:谁提出了创意?我们引入IdeaLens,一个检测器,用于识别文档的创意是来自人类还是AI(创意来源),无论谁写了文字。为了使IdeaLens聚焦于创意而非文风,我们将文档表示为大纲:项目列表,每项将话语角色与内容的简短释义描述配对,最小化与原始文本的词汇重叠。我们在100万份FineWeb文档上训练IdeaLens,使用来自Pangram(一种文风来源检测器)的银标签。由于大纲在很大程度上剥离了表面信息,标签必须主要通过创意来拟合。在一项对照研究中,随着模型根据越来越详细的人类计划写作,IdeaLens的AI标记率从95%降至7%,而Pangram 4仍标记92%;对于AI衍生的计划,IdeaLens保持在96%以上。相反,在一个由人类作者根据AI生成计划撰写的50个故事的新数据集上,IdeaLens将68%的故事标记为AI,而Pangram 4为8%。在包含19个现有检测基准的全面套件上,我们表明IdeaLens在低误报率下保持强检测率,表明创意本身提供了强大的判别信号,其性能在领域、格式和语言上保持一致。最后,我们检查了IdeaLens的9万条预测,以表征人类与AI创意之间的系统性差异。我们发布我们的模型和标注数据集,以促进未来创意来源检测的研究。
英文摘要
While modern AI detectors identify who wrote the words, emerging policies on AI use increasingly hinge on a different question: who came up with the ideas? We introduce IdeaLens, a detector that identifies whether a document's ideas came from a human or AI (idea provenance), regardless of who wrote its words. To focus IdeaLens on ideas rather than prose, we represent documents as outlines: lists of items that each pair a discourse role with a brief, paraphrased description of the content, minimizing word-level overlap with the raw text. We train IdeaLens on 1M FineWeb documents with silver labels from Pangram, a prose provenance detector. Since the outlines are largely stripped of surface-level information, the labels must be fit mainly through the ideas. In a controlled study, IdeaLens's AI flag rate drops from 95% to 7% as models write from increasingly detailed human plans, while Pangram 4 still flags 92%; from AI-derived plans, IdeaLens stays above 96%. Conversely, on a new dataset of 50 stories that human authors wrote from AI-generated plans, IdeaLens flags 68% of the stories as AI, compared to 8% for Pangram 4. On a comprehensive suite of 19 existing detection benchmarks, we show that IdeaLens maintains strong detection rates at low false positive rates, suggesting that ideas themselves provide a powerful discriminative signal, and its performance holds across domains, formats, and languages. Finally, we examine 90K predictions from IdeaLens to characterize systematic differences between human and AI ideation. We release our models and labeled datasets to facilitate future research on idea provenance detection.
Comments53 pages (9 main), 7 figures, 50 tables. Code: https://github.com/RishanthRajendhran/IdeaLens Models and data: https://huggingface.co/collections/rishanthrajendhran/idealens-6abee785ce6196fc0be9200f Demo: http://ideadetector.ai/