发表机构
Polytechnic University of Marche; University of Modena and Reggio Emilia(马尔凯理工大学; 摩德纳大学和雷焦艾米利亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对人工智能生成文本检测问题,提出几何轨迹和对比学习框架GTCL,通过分割文档为局部单元、编码并构建序列级表示,再应用对比学习学习几何规律,实验表明此方法优于基线,为AIGTD开辟新动态方向。
AI 中文摘要
大多数现有的人工智能生成文本检测(AIGTD)方法将文档视为静态对象,并基于聚合统计或全局压缩嵌入做出决策。然而,这种观点忽略了自回归生成的内在动态性质,即内容通过潜在空间逐步演变。本文将AIGTD重新表述为区分潜在生成轨迹的问题。我们不依赖静态表示,而是对文本表示如何在序列中演变进行建模。为此,我们提出了几何轨迹和对比学习(GTCL)框架,该框架将文档分割成有序的局部单元,在嵌入空间中对每个单元进行编码,并构建结构化的序列级表示。然后,GTCL对这些轨迹应用对比学习,以学习与自回归生成相关的几何规律。在三个不同基准上的评估表明,GTCL始终优于检测基线,这意味着显式建模序列动态可以在不同模型和领域中提供强大的判别信号。这些结果表明,对轨迹差异进行建模可以改进检测,并开辟一个在以前的AIGTD文献中未充分探索的动态方向。
英文摘要
Most existing approaches to AI-Generated Text Detection (AIGTD) treat documents as static objects and base their decisions on aggregate statistics or globally compressed embeddings. However, this perspective overlooks the inherently dynamic nature of autoregressive generation, where content evolves progressively through the latent space. In this paper, we reformulate AIGTD as the problem of distinguishing between latent generation trajectories. Instead of relying on static representations, we model how textual representations evolve across the sequence. To this end, we propose Geometric Trajectory and Contrastive Learning (GTCL), a framework that segments the document into ordered local units, encodes each unit in an embedding space, and constructs a structured and sequence-level representation. GTCL then applies contrastive learning to these trajectories to learn geometric regularities associated with the autoregressive generation. Evaluations performed on three different benchmarks and several approaches show that GTCL outperforms detection baselines consistently, which implies that explicitly modeling sequential dynamics provides robust discriminative signals across models and domains. These results suggest that modeling trajectory differences could improve detection and open up a dynamic direction that has been underexplored in previous AIGTD literature.