发表机构
Stanford University; University of Pennsylvania; Cornell University(斯坦福大学; 宾夕法尼亚大学; 康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究LLM检测对下游指标的影响,开发程式化模型捕捉用户策略行为。发现不完善的检测器会扭曲LLM使用及输出质量,导致用户增加使用量且输出质量可能降低,还呈现“先升后降”检测模式,揭示了检测干预的失效模式。
AI 中文摘要
随着大语言模型(LLM)的应用越来越广泛,人们对检测LLM生成的内容兴趣日增,例如通过检测工具和基于语言模式的启发式方法。检测器作为一种干预手段,不仅影响被检测的属性本身,还影响诸如LLM使用情况和输出质量等下游指标。本文展示了不完善的LLM检测器如何通过扭曲用户在工作流程中使用LLM的动机,对这些下游指标产生违反直觉的影响。我们开发了一个程式化模型,捕捉用户如何策略性地选择使用LLM的程度以及如何对内容进行后处理以减少被检测属性。通过该模型表明,LLM检测可能会反常地导致用户增加LLM使用量。此外,即使减少被检测属性会提高输出质量,引入LLM检测器也可能导致用户产生质量更低的输出。相比之下,检测器会使被检测属性呈现出清晰的“先上升后下降”模式,我们在arXiv摘要的词频上通过实验重现了这一模式。总之,我们的工作说明了LLM检测如何扭曲LLM使用和输出质量,揭示了LLM检测器作为对这些下游指标的干预手段时的失效模式。
英文摘要
As LLM adoption becomes more widespread, there is a growing interest in detecting LLM-generated content, for example through LLM detection tools and through heuristics based on language patterns. Detectors operate as an intervention that steers not only the detected attribute itself, but also downstream metrics such as LLM usage and output quality. In this work, we demonstrate how imperfect LLM detectors lead to counterintuitive impacts on these downstream metrics, by distorting how users are incentivized to use LLMs in their workflow. We develop a stylized model which captures how users strategically choose how much to use the LLM and how to post-process content to reduce the detected attribute. Using this model, we show that LLM detection can counterintuitively lead humans to increase their LLM usage. Moreover, even when reducing the detected attribute improves output quality, we find that introducing an LLM detector can lead users to produce lower quality outputs. In contrast, we show that detectors result in a clean "rise-then-fall" pattern for the detected attribute, which we empirically reproduce for word frequencies on arXiv abstracts. Altogether, our work illustrates how LLM detection can distort LLM usage and output quality, uncovering failure modes when LLM detectors operate as an intervention on these downstream metrics.