发表机构
University of North Texas(北德克萨斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究发现大语言模型审计文档时,批量增大导致检测崩溃,从单文档50%降至大批量2.8%,且失败模式为自信捏造而非弃权,需有界批量、直接注入和机械验证。
AI 中文摘要
大型语言模型日益被提议作为文档质量的自动化审计者,但其作为植入错误检测器的可靠性尚未得到充分表征。我们构建了一个包含150篇学术论文的污染语料库,涵盖供应链管理和医学研究领域,注入了450个已知污染物,分为三种类型:拼写错误、语义反转和荒谬的脱离上下文插入。随后,我们评估了Google Gemini 3.0 Pro在三种规模递增的提示机制下(单文档、小批量和大批量)从60篇文档中恢复包含180个污染物的答案键子集的能力。检测在小规模下保持,随后崩溃:单文档恢复率为50%,小批量恢复率为60%,大批量恢复率仅为2.8%。大规模下的失败模式并非弃权(不执行),而是捏造。模型并未报告处理不完整,而是产生了自信的发现,包括其自身发明的污染物,例如“心灵感应松鼠”和“量子动力烤面包机”等荒谬内容,这些模仿了植入材料的风格,但未出现在任何文档中。检测也因污染类型而异:在完成的评估中,荒谬插入的恢复率为75%,而语义反转和拼写错误的恢复率各仅为50%。最可能在现实中发生的污染,即看似合理的污染,往往最容易被遗漏。我们得出结论:大语言模型文档审计的退化并非优雅地发生,而是具有欺骗性,并概述了此类系统所需的约束机制:有界的批量大小、直接内容注入,以及对每项报告发现与源文本进行机械验证。
英文摘要
Large language models are increasingly proposed as automated auditors of document quality, yet their reliability as detectors of planted errors is poorly characterised. We construct a contaminated corpus of 150 academic papers spanning supply chain management and medical research, injecting 450 known contaminants of three types: typographical corruption, semantic reversal, and absurd out-of-context insertion. We then evaluate Google Gemini 3.0 Pro's ability to recover a 180-contaminant answer-key subset across 60 documents under three prompting regimes of increasing scale: single document, small batch, and large batch. Detection is unreliable even at small scale and collapses entirely at large scale: 50% recovery on single documents and 60% on small batches, a difference this sample cannot resolve, against 2.8% on large batches. The failure mode at scale is not abstention but fabrication. Rather than reporting incomplete processing, the model produced confident findings including invented contaminants of its own, absurdities such as "telepathic squirrel" and "quantum-powered toaster" that mimic the style of the planted material but do not appear in any document. Detection also varies by contamination type: absurd insertions were recovered at 75% in completed evaluations, while semantic reversals and typographical corruptions were each recovered at only 50%. The corruptions most likely to occur in the wild, plausible ones, are the ones most often missed. We conclude that LLM document auditing degrades not gracefully but deceptively, and outline the harness such systems require: bounded batch sizes, direct content injection, and mechanical verification of every reported finding against source text.
Comments8 pages, 3 tables. v2 removes an evaluation-scoring claim that was not independently observed, adds confidence intervals and significance tests, and states the batch size and retrieval confound explicitly. Preprint also deposited at Zenodo, doi:10.5281/zenodo.21939087