面向在线调查中AI辅助回答的检测
Towards Detecting AI-Assisted Responses in Online Surveys
浏览论文内容
中文总结 AI 辅助
本研究提出ASURRE基准数据集,用于检测在线调查中的AI辅助回答,发现基于人设的智能体难以检测,但可通过行为线索聚合将AUROC提升0.14。
中文摘要 AI 辅助
使用大语言模型(LLM)完成在线调查会损害基于调查的研究的有效性,但检测此类使用情况仍未被充分探索。我们引入了一个初始基准数据集,即ASURRE,用于AI辅助调查参与,以捕捉从完全生成、修订到基于人设的智能体式完成等使用策略。在这些策略的控制下,我们使用多个LLM在三个不同学科的真实世界调查上生成了LLM辅助的调查回答,并与真实的人类回答配对。我们对现有机器生成文本(MGT)检测器的评估表明,天真的AI使用易于检测,而模仿整个受访者的基于人设的智能体则将检测器性能推向随机水平。我们进一步表明,智能体式完成无法完全复制受访者层面的行为,并留下独特的行为痕迹。虽然单个线索可以通过有针对性的提示来规避,但一个简单的、基于少样本的、无需训练的线索聚合器在智能体设置下将平均AUROC比最佳现有检测器提高了+0.14。我们的项目可在以下https URL获取。
英文摘要
The use of LLMs to complete online surveys impacts the validity of survey-based research, but detecting such usage remains underexplored. We introduce an initial benchmark dataset, namely ASURRE, for AI-assisted survey participation to capture usage strategies ranging from full generation and revision to persona-grounded agentic completion. Controlled by these strategies, LLM-assisted survey responses are generated using multiple LLMs on three real-world surveys in different disciplines, paired with genuine human responses. Our evaluation of existing machine-generated text (MGT) detectors shows that naive AI usage is readily detectable, whereas persona-grounded agents that mimic entire respondents push detector performance toward chance. We further show that agentic completion cannot fully replicate respondent-level behaviour and leaves distinctive behavioural traces. While individual cues can be circumvented by targeted prompting, a simple few-shot, training-free aggregator over these cues improves mean AUROC by +0.14 over the best existing detector across agentic settings. Our project is available at https://github.com/mike-qz-wang/ASURRE.
发表机构
- The University of Melbourne(墨尔本大学)
- Deakin University(迪肯大学)
机构由 AI 辅助整理,请以论文原文为准。