arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31054cs.AI

廉价、开放智能体使大语言模型污染更难缓解

Cheap, open agents make LLM pollution harder to mitigate

Raluca Rilla, Anne-Marie Nussberger, Rui Mata, Dirk U. Wulff

首次发表
浏览论文内容

中文总结 AI 辅助

本研究比较九种智能体配置,发现完全开放的智能体成本低且性能与商业相当,但更难被检测,构成LLM污染的独特风险,需多层检测策略,尤其重视开放文本分析。

中文摘要 AI 辅助

大语言模型(LLM)污染发生在合成响应污染旨在捕捉人类行为的数据时。高昂的部署成本迄今限制了自主调查智能体带来的风险。然而,开放权重模型与开源智能体框架的结合可能已消除这一障碍。我们比较了九种智能体配置的性能和可检测性,范围从完全开放的变体到封闭的商业变体。每个智能体自主完成了一项包含多种响应类型的调查,从而产生各种检测检查。完全开放的智能体在本地运行,无需使用费用,且性能与商业替代方案相当。开放和商业智能体未能通过不同的检查集,没有单一检查能可靠地检测所有智能体,但开放文本响应在区分智能体和人类方面效果最佳。这些发现将完全开放的智能体确定为LLM污染的独特风险,并支持强调开放文本分析的多层检测策略。

英文摘要

Large Language Model (LLM) pollution occurs when synthetic responses contaminate data intended to capture human behavior. High deployment costs have so far limited the risk posed by autonomous survey agents. However, open-weight models paired with open-source agentic frameworks may have removed this barrier. We compared the performance and detectability of nine agent configurations, ranging from fully open variants to closed commercial ones. Each agent autonomously completed a survey containing multiple response types yielding various detection checks. Fully open agents ran locally without usage fees and performed competitively with commercial alternatives. Open and commercial agents failed different sets of checks, and no single check reliably detected all agents, but open-text responses discriminated best between agents and humans. These findings identify fully open agents as a distinct risk for LLM pollution and support multilayered detection strategies emphasizing open-text analysis.

发表机构

  • Max Planck Institute for Human Development(马克斯·普朗克人类发展研究所)
  • International Max Planck Research School on Learning, Institutions, and Future Evolution (LIFE)(国际马克斯·普朗克学习、制度与未来演化研究学院)
  • University of Basel(巴塞尔大学)
  • Vienna University of Economics and Business(维也纳经济与商业大学)

机构由 AI 辅助整理,请以论文原文为准。

↑