arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MD-ProTector:为大语言模型生成文本检测定位多个数据驱动原型

MD-ProTector: Positioning Multiple Data-Driven Prototypes for LLM-Generated Text Detection

Jinmo Han, Jimin Hong, Chanyeong Moon, Ju Yeon Kang, Seonuk Kim, Nam Soo Kim

arXiv 2608.10459首次发表:更新:

发表机构

Seoul National University(首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对LLM生成文本检测的类别内部差异问题,提出MD-ProTector模型,通过原型定位损失优化,在多基准测试中取得优于同类编码器方法的检测性能。

AI 中文摘要

随着大语言模型(LLM)生成内容愈发复杂,用于区分这类文本与人类撰写文本的检测系统必须能够规模化运行,同时处理多样的写作风格、领域、语言以及生成器模型。仅输入编码器检测器适用于实际部署场景,但标准二分类仅提供类别标签,未明确组织两类文本内部的大量差异。我们提出MD-ProTector,其在编码器嵌入空间中用多个可训练参考向量(称为原型)表示每个类别,这些原型为同一类别内不同文本组提供独立决策边界。然而,仅添加多个原型无法确定每个原型应代表哪种差异,MD-ProTector通过原型定位损失解决该问题,该损失将类别级结构与区分单个原型的类别内差异分离。在涵盖领域、生成器、语言和对抗性差异的三个大规模基准的五个设置中评估,MD-ProTector在MAGE CDCM和RAID上实现最高平均召回率(AvgRec),在RAID上实现最高受试者工作特征曲线下面积(AUROC)和最低95%假阳性率(FPR95),优于所有对比的基于编码器的方法。

英文摘要

As LLM-generated content becomes more sophisticated, detection systems for distinguishing those texts from human-written text must operate at scale while handling diverse writing styles, domains, languages, and generator models. Input-only encoder detectors are suitable for practical deployment setting, but standard binary classification supplies only the class label and does not explicitly organize the substantial variation within either class. We propose MD-ProTector, which represents each class with multiple trainable reference vectors in the encoder embedding space, referred to as prototypes. These prototypes provide separate decision boundaries for different groups of texts within the same class. However, adding multiple prototypes alone does not determine which variation each prototype should represent. MD-ProTector addresses this problem with Prototype Positioning loss, which separates class-level structure from the within-class variation that differentiates individual prototypes. Evaluated across five settings from three large-scale benchmarks covering domain, generator, language, and adversarial variation, MD-ProTector achieves the highest AvgRec on MAGE CDCM and RAID and the highest AUROC and lowest FPR95 on RAID among the compared encoder-based methods.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑