发表机构
Institute for Complex Social Dynamics, Carnegie Mellon University; Carnegie Mellon University(卡内基梅隆大学复杂社会动力学研究所; 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究探讨人工智能时代科学欺诈新形式,即间接数据投毒。对手 corrupt 开放数据集上传,使自主研究代理处理后传播欺诈。通过多话题、多系统实验发现投毒成功率高而检测率低。提出科学家角色和数据溯源审计措施,后者可有效降低攻击成功率。
AI 中文摘要
科学欺诈是恶意实体在科学领域制造争议的手段。过去需要公司资源,如今人工智能日益自动化科研,我们探讨远程对手能否利用人工智能的正当使用损害科学诚信。设想并实证评估间接数据投毒攻击,对手 corrupt 开放数据集并上传中毒变体到公共库,自主研究代理可能处理该数据,使诚实科学家成为欺诈传播者。在五个社会热点话题、三个前沿人工智能系统及 450 次实验中,投毒成功率达 49.56%,检测率仅 6.0%。攻击无需特定触发词等,仅需开放数据生态系统和误导性元数据。为缓解攻击,提出并评估科学家角色和数据溯源审计两种措施。科学家角色仍使 16.67%的实验得出中毒结论,溯源审计可将攻击成功率降至零。结果表明间接数据投毒可能使科学欺诈达到前所未有的规模,但可通过数据检索时的适当审计缓解。
英文摘要
Scientific fraud is the instrument of doubt that malicious entities can use to establish controversy in science. Historically, it required the resources of a company: deep pockets, ghostwritten articles, and corrupt academics. Today, Artificial Intelligence (AI) is increasingly automating scientific research, so we ask: Can a remote adversary weaponize the honest use of AI in science to compromise scientific integrity? We envision and empirically evaluate a new attack, indirect data poisoning, in which an adversary corrupts an open dataset and uploads the poisoned variant to a public repository. Autonomous research agents may independently retrieve and process this data, turning honest scientists into the unpaid and unwitting distributors of fraud at scale. Across five socially-salient topics, from hiring discrimination to the safety of autonomous vehicles, three widely used frontier AI systems (Claude Code with Claude Opus 4.7, Codex with GPT-5.5, Gemini CLI with Gemini 3.1 Pro), and 450 ethically contained experimental runs, we find that poisoning succeeds in 49.56% of runs, while the rate of poisoning detection is only 6.0%. The attack requires no topic-specific trigger-words, agent access, indirect prompt injection, or fabricated papers, only the open data ecosystem and misleading metadata. To mitigate the attacks, we propose and evaluate two measures: a scientist persona and a data provenance audit with five checks (referencing papers, social markers, statistical anomalies, related datasets, poisoning caution). We find that the persona still leaves 16.67% of runs with a poisoned conclusion, but provenance auditing reduces attack success rate to zero. Our results suggest that indirect data poisoning may enable scientific fraud at unprecedented scale, but these attacks can be mitigated with suitable auditing by agents during data retrieval.