arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.29539cs.CLcs.AI

ARB:用于AI文本检测器评估的匹配作者改写基准数据集

ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation

Gaetano Perrone, Simon Pietro Romano

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出用于AI文本检测器评估的ARB基准数据集,对比五种检测器在传统基准与人类被LLM改写场景下的性能,发现传统基准测得的性能无法迁移至改写人类文本场景。

中文摘要 AI 辅助

现有的AI文本检测基准通常将人类撰写的文本与大型语言模型(LLM)直接生成的文本进行对比。虽然已有研究表明改写和 paraphrasing(意译)会降低检测器的性能,但目前仍不清楚在这种传统基准上测得的性能,能否预测当人类撰写的内容被LLM改写时检测器的表现。为解决这一差距,本文引入了作者改写基准(Authorship-Rewriting Benchmark,ARB),该数据集由1800份人类源文本构建而成,分别来自XSum、WritingPrompts和OpenWebText各600份,同时包含四个开源权重生成器:Llama-3.2-3B、Qwen2.5-7B、Mistral-7B、Gemma-2-9B。每份源文本生成四种匹配变体:人类撰写文本(HUMAN)、LLM直接生成文本(Free-LLM)、LLM改写的人类文本(H2L)、同一生成器改写的LLM文本(LLM2L)。本文在严格的1%误报率(TPR@1%FPR)操作点下评估了五种检测器:FastDetectGPT、Binoculars-falcon-7b、RADAR、BERT-Defense、RoBERTa-Defense。结果显示,FastDetectGPT和Binoculars-falcon-7b能检测出91.2%和93.5%的LLM直接生成文本,但仅能检测出30.8%和15.1%的LLM改写人类文本,降幅达60至78个百分点;当LLM文本被同一模型改写时,上述检测器的召回率仍保持78.3%和83.0%,降幅仅为10至13个百分点,RADAR也呈现相同模式(从66.8%降至12.2%),而BERT-Defense和RoBERTa-Defense在所有场景下的召回率均低于3%。这些结果表明,在传统的人类与LLM对比基准上测得的检测器性能,无法迁移到被LLM改写的人类文本场景中,不过同一检测器对仅由LLM改写的文本仍基本保持鲁棒性。

英文摘要

Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs). While prior work has shown that rewriting and paraphrasing can degrade detector performance, it remains unclear whether performance measured on this conventional benchmark predicts detector behavior when human-authored content is rewritten by an LLM. To address this gap, we introduce Authorship-Rewriting Benchmark (ARB), built from 1,800 human source texts (600 each from XSum, WritingPrompts, and OpenWebText) and four open-weight generators (Llama-3.2-3B, Qwen2.5-7B, Mistral-7B, Gemma-2-9B). Each source item yields four matched variants: human-written (HUMAN), direct LLM generation (Free-LLM), LLM-rewritten human text (H2L), and same-generator LLM-rewritten LLM text (LLM2L). We evaluated five detectors (FastDetectGPT, Binoculars-falcon-7b, RADAR, BERT-Defense, RoBERTa-Defense) at a strict 1%-false-positive operating point (TPR@1%FPR). FastDetectGPT and Binoculars-falcon-7b detected 91.2% and 93.5\% of direct LLM text, but only 30.8% and 15.1% of human text an LLM had rewritten, a drop of 60-78 percentage points. The same detectors retained 78.3% and 83.0% recall when LLM text was rewritten by the same model, a much smaller decline of 10-13 points. RADAR followed the same pattern (66.8% to 12.2%), while BERT-Defense and RoBERTa-Defense stayed below 3% recall across all regimes. These results show that detector performance measured on the conventional human-vs-LLM benchmark does not transfer to human-authored text revised by an LLM, even though the same detectors remain largely robust to LLM-only rewriting.

发表机构

  • University of Napoli Federico II(那不勒斯费德里科二世大学)

机构由 AI 辅助整理,请以论文原文为准。

↑