告诉我你如何推理,我就能说出你是谁:用于可靠大语言模型作者归属的推理图
Show Me How You Reason and I'll Tell You Who You Are: Reasoning Graphs for Robust LLM Authorship Attribution
浏览论文内容
中文总结 AI 辅助
针对LLM生成文本检测和作者归属问题,此前方法易受改写等影响。本文提出利用论证挖掘管道提取推理图的图神经网络方法,在混淆攻击及新模型版本文本评估中,相比传统基线展现更强鲁棒性和泛化能力,提升了检测准确率。
中文摘要 AI 辅助
鉴于当前在几乎任何可想象的场景中使用大语言模型(LLM)的趋势,LLM生成文本的检测和作者归属已成为一个紧迫问题。先前工作主要集中在表面语言特征,易受改写和其他混淆技术影响。本文超越语言表面,提取并分析LLM生成文本中的推理结构,以捕捉更复杂的作者归属信号。我们提出一种图神经网络方法,利用论证挖掘管道提取的推理图,相较于传统Longformer基线展现出更高的鲁棒性和泛化能力。在改写和回译等混淆攻击下,我们的方法比基线高出27个百分点;在评估由未见模型版本生成的文本时,高出19个百分点,模拟了新LLM版本不断发布的现实情况。
英文摘要
Given the current trend to employ large language models (LLMs) in almost any imaginable context, LLM-generated text detection and authorship attribution have become a pressing issue. Prior work has primarily focused on surface-level linguistic features, an approach shown to be susceptible to paraphrasing and other obfuscation techniques. In this paper, we go beyond the linguistic surface, extracting and analysing reasoning structures in LLM-generated texts with the goal of capturing more complex signals of LLM authorship. We propose a graph neural network approach that leverages reasoning graphs extracted by an argument mining pipeline, demonstrating improved robustness and generalisation over a traditional Longformer baseline. Our approach outperforms the baseline by up to 27 percentage points under the obfuscation attacks such as paraphrasing and backtranslation, and 19 percentage points when evaluated on the texts generated by the unseen model versions, simulating real-world conditions in which new LLM versions are continuously released.
发表机构
- University of Passau(帕绍大学)
- University of Dundee(邓迪大学)
机构由 AI 辅助整理,请以论文原文为准。